Capability benchmark · Materials science · Computer science

2D-FET-Bench finds expert acceptance well below automated layout success

128 executable flake-specific transistor-layout tasks and an expert audit of 436 verifier-passing layouts.

Summary

The six-model panel’s strongest configuration, GPT5.6-Luna ReAct-3, passes 62.3% of attempts. A human expert accepts 260 of 436 sampled verifier-passing layouts across five configurations (59.6%). Each sample is one passing layout per covered task; 128 scripted references pass the geometric checks.

AI role

Agents translate textual device specifications and microscopy-derived contour coordinates into editable GDSII layout geometry.

Narrative role

The audit distinguishes executable geometric success from device-level acceptability in an actual materials-research design workflow.

Caveat

Inputs are coordinates rather than raw images; no fabrication or electrical testing. One expert reviews different covered sets across configurations. The manual-time comparator is an estimate, and public code/data release is promised rather than independently replayed.