Skip to content

Some questions  #7

Description

@abdelkareemkobo

Thanks for creating and sharing this tool

I am native Arabic and u found this tool has a great potential.
I have quick questions.

  1. How can you handle the process of translate multiple samples at the same time knowing that every sample will have different size one maybe 500 tokens and the other is 1000 tokens... For Arabic i didn't find a good models from translation expect the cohere models and i am taking about models in sizes from 4b to 12b that i can try.
    I found theses models are very good even with technical ones but they can't handle multiple sentences and you must play a lot with the generation and specify the maximum number of tokens or it will continue generating.
  2. For splitting, did you tried chonkie? The Neural chuker i think will help you a lot.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions