Phi-1 and the Small Code Model Trained on ‘Textbook Quality’ Data
Microsoft’s phi-1 showed that a 1.3B-parameter code model trained on curated and synthetic “textbook quality” data could achieve strong results on compact Python coding benchmarks.
Microsoft’s phi-1 showed that a 1.3B-parameter code model trained on curated and synthetic “textbook quality” data could achieve strong results on compact Python coding benchmarks.
WizardCoder adapted Evol-Instruct to programming, using synthetically evolved coding instructions to turn pretrained code models into stronger natural-language task followers.
PolyCoder released a 2.7-billion-parameter multilingual code model in 2022, giving researchers open weights for studying code generation while many leading systems remained closed.
SantaCoder let the BigCode collaboration test governed training data, filtering, fill-in-the-middle modeling, and responsible model licensing before scaling the approach into StarCoder.