Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

There’s typically a difference in LR between a ‘continued pretrain’ and ‘fine tune.’ I don’t have the details around miqu, but was merely trying to say that Mistral could produce a better version of these models than the OSS community might. If the size of the corpora they use means we are no longer in fine tuning territory, then okay.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: