Capability benchmark · Computer science · General science
SciCode evaluates models on scientist-curated research coding
Model performance on realistic scientific coding tasks across natural-science subfields.
Summary
SciCode provides scientist-curated coding tasks that are closer to research practice than ordinary programming puzzles.
AI role
Generated code for scientific problems requiring domain context, formulas, dependencies, and multi-step reasoning.
Narrative role
This is a supporting capability signal for whether AI systems can handle the computational substrate of modern science.
Caveat
Scientific coding success does not directly imply discovery, especially when experimental design and interpretation are outside the benchmark.