Capability benchmark · Computer science · General science

SciCode evaluates models on scientist-curated research coding

Model performance on realistic scientific coding tasks across natural-science subfields.

Summary

SciCode provides scientist-curated coding tasks that are closer to research practice than ordinary programming puzzles.

AI role

Generated code for scientific problems requiring domain context, formulas, dependencies, and multi-step reasoning.

Narrative role

This is a supporting capability signal for whether AI systems can handle the computational substrate of modern science.

Caveat

Scientific coding success does not directly imply discovery, especially when experimental design and interpretation are outside the benchmark.