Hi team, multilingual performance looks strong in my tests! Are there plans to adapt MonkeyOCRv2 for visual document retrieval, video scene text tracking, or extend handwriting recognition evaluations to datasets?
It would be interesting to see how the vision backbone transfers to these settings. Thanks!
Hi team, multilingual performance looks strong in my tests! Are there plans to adapt MonkeyOCRv2 for visual document retrieval, video scene text tracking, or extend handwriting recognition evaluations to datasets?
It would be interesting to see how the vision backbone transfers to these settings. Thanks!