BigCode, a joint initiative of Hugging Face and ServiceNow, has introduced StarCoder and StarCoderBase, two open-source models for working with code. The project pairs broad language support with an emphasis on how training data is selected and how developers can check whether their work was included.
Models built to handle more code at once
The StarCoder models have 15.5 billion parameters and can generate code in 86 programming languages. Their 8K context windows let them process larger amounts of code, supporting both code understanding and generation.
The researchers used a technique called “multi-query attention.” The article describes this as allowing the models to focus on multiple parts of code at once rather than processing each token in turn. The stated result is faster, more efficient handling of longer code contexts.
That combination could matter when a task depends on understanding relationships across a larger code sample. A model that can take in more code at once has more context for generating or interpreting a response, although the article does not claim that context alone guarantees correct output.
Curated training data and model performance
Training data selection involved substantial manual work, according to participating researcher Lubna Ben Allal. The team inspected 50–100 files for all the extensions in the selected programming languages and chose filters they considered appropriate.
The article reports that both models performed better in benchmarks than other open models supporting multiple programming languages. It also says they equaled or surpassed OpenAI’s “code-cushman-001” model in some comparisons. On HumanEval, StarCoder outperformed that model and all open code generation models; the reported gap was larger on DS-1000.
DS-1000 covers more diverse and realistic data science problems spanning 7 libraries. That makes it a different kind of comparison from a benchmark focused on code generation more generally. The results position StarCoder as a strong open model in the evaluations described, while the article notes that its code performance may still lag GPT-4.
Data governance is part of the release
BigCode says it used permissible data without personal references for training. The team also provided an opt-out mechanism and a code snippet search engine so developers can check whether their code is included in data from The Stack database.
These measures give developers ways to inspect potential inclusion and request an opt-out. They also make data selection a visible part of the project, rather than treating it as a background detail of model development.
What developers can do with StarCoder
The models are released under the Open Responsible AI Model license, which supports commercial use. They are not instruction-optimized out of the box, but the article says they can be optimized into a technical assistant with additional instructions.
For developers, the release combines multilingual code generation, a longer context window, benchmark results, and stated data governance measures. The models may be useful to explore as open alternatives, but their intended use as a technical assistant requires additional instruction work. Further information and links are available through Hugging Face StarCoder.