Your previous report included a single task SFT stage in the training pipeline, while also noting that SFT could be counterproductive for certain tasks. I would like to clarify which specific tasks—or tasks with what characteristics—exhibited this negative impact?
Due to limited hardware resources, I am restricted to using LoRA for fine-tuning. My goal is to see a rapid improvement in success rates starting from your pre-trained base model(Q 0.19). Could you recommend which types of tasks I should prioritize for initial experiments, and what specific configurations you would suggest?"
Your previous report included a single task SFT stage in the training pipeline, while also noting that SFT could be counterproductive for certain tasks. I would like to clarify which specific tasks—or tasks with what characteristics—exhibited this negative impact?
Due to limited hardware resources, I am restricted to using LoRA for fine-tuning. My goal is to see a rapid improvement in success rates starting from your pre-trained base model(Q 0.19). Could you recommend which types of tasks I should prioritize for initial experiments, and what specific configurations you would suggest?"