Skip to content

Size of data for continued pretraining #7

Description

@AmgadHasan

Hello!

Thank you so much for developing and releasing this model to the public. As a native Arabic speaker, I highly appreciate your efforts in enriching our beautiful language.

I have the following question related to the training process:

As per my understanding, the first step is continued pretending of Llama2 on Arabic data in a self supervised manner.
My question is how big is the data used in this step?

Thanks in advance.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions