Evaluation · official scores, engineering progress + IC sign-off

Measured, not claimed.

Official benchmark scores use clean-room blind runs and each benchmark's official scorer. Engineering checkpoints are reported separately with their scope, exposure history and unresolved cases. The IC results document end-to-end spec → GDSII work through open-source sign-off; each result retains its evidence and limitations.

Open Benchmarks · fully blind · official upstream testbenches

Open-benchmark scores

Every open-benchmark score below follows clean-room blind rules: the authoring agent sees only the permitted prompt/context — never the hidden testbench, golden, or sibling solutions — and the benchmark's official scorer establishes the result. Each cell names the AI final semantic authority and Vibe-IC plugin version. The newest completed results are the 2026-09-06 single-shot blind runs on shipped v1.17.60 (plugin tree 1eb2a241) with Claude Opus 5 as final reviewer for VerilogEval-v2 (152/156) and VerilogEval-Human (154/156, tree a938bc0a), and Claude Fable 5 on shipped v1.15.20 for RTLLM (48/50, 2026-09-01). Beside each VerilogEval score we state the theoretical maximum: problems proven broken with re-runnable evidence — a golden that fails its own testbench, or contradicts its own prompt — stay in the denominator and are named, never rewritten. All evidence cells are public.

97.44%Claude Opus 5 · shipped v1.17.60 · VerilogEval-v2 — 152/156 pass@1 = theoretical maximum (4 problems proven broken)
98.72%Claude Opus 5 · tree a938bc0a · VerilogEval-Human — 154/156 pass@1 = theoretical maximum (2 problems proven broken)
96%Claude Fable 5 · v1.15.20 · RTLLM — official 48/50 pass@1; 46/48 on discriminating testbenches
4/4Open ICs spec→GDSII — PASS_WITH_WAIVERS (strict)

Measured datapoints · model + plugin identified

Authoring-model scoreboard

All rows use the canonical clean-room entry and official scorer, but they were not run on one identical plugin build. Read each cell as an AI-authority + plugin datapoint, not as a controlled model-only A/B. Single-shot pass@1 and terminal loop2converge are never merged. The 2026-08-31 VE cells started from v1.13.78 plus four disclosed campaign paths, so they are not shipped-release reproductions; their complete public evidence is linked below.

Benchmark · blind · official scorer Claude GPT Kimi
VerilogEval-v2 (156) 152/156 pass@1 = max 152
Claude Opus 5 · shipped v1.17.60 · 2026-09-06
151/156 pass@1 → 153/156 terminal
GPT-5.6-Sol · campaign-patched v1.13.78
147/156 single
Kimi K3 · plugin version not bound
VerilogEval-Human (156) 154/156 pass@1 = max 154
Claude Opus 5 · tree a938bc0a · 2026-09-06
152/156 pass@1 → 153/156 terminal
GPT-5.6-Sol · campaign-patched v1.13.78
149/156 single
Kimi K3 · plugin version not bound
RTLLM v2.0 (50) 48/50 pass@1
Claude Fable 5 · v1.15.20 · 46/48 discriminating
48/50 pass@1
GPT-5.6-Sol · v1.14.5 · 45/47 discriminating
Not measured

Read-out: the 2026-09-06 cells on shipped v1.17.60 are single-shot blind runs. VerilogEval-v2 records 152/156, which is its theoretical maximum: four problems are proven broken with re-runnable evidence in each cell's theoretical_max.json — Prob099 (the golden cannot compile against its own testbench), Prob062, Prob093 and Prob149 (the golden contradicts what the prompt states). VerilogEval-Human records 154/156, its theoretical maximum (Prob062 and Prob093). On RTLLM the newest published cell (Claude Fable 5 · shipped v1.15.20, 2026-09-01) records 48/50 official pass@1 (46/48 discriminating), matching GPT-5.6-Sol · v1.14.5's 48/50; its pre-scorer Program First + AI review layer accepted 50/50, and that semantic layer does not rewrite the official score. Its two residuals are one proven dataset defect (the radix2_div golden fails its own testbench, 3/8) and one plugin-gate misdirection captured as public issue vibe-ic#1998. The earlier 153/156 on VerilogEval-v2 included Prob149, which was passable only through a lessons entry later found to be oracle-derived and withdrawn. Denominators are always the original ones.

Plus: CVDP 243/302 = 80.46% official-compliant blind pass@1 on the public no-commercial code-generation set (measured on Opus 4.8 · plugin v1.2.63 — a separate earlier campaign) — prompt+context-only, with the hidden cocotb harness and golden kept off-limits; and a MetRex premise-dissolving demo — instead of predicting post-synthesis area, Vibe-IC runs MetRex's exact Yosys + sky130 recipe and reproduces its golden area to a 1.15% median (90% within 5%). Honest caveats included: dataset, specification, and scorer/tool-dialect limitations are recorded with the exact official score, never silently dropped.

Engineering progress · 2026-09-06 · original unfinished69 subset

CVDP hard-tail engineering checkpoint

Engineering acceptance is 68/69 on the original unfinished69 hard-tail subset, with 1 case unresolved. This checkpoint covers that fixed subset only. It is not a full-dataset score, Pass@1, or an all-pass result.

Official score: NOT_RUN. Blindness status: NOT_ELIGIBLE_AS_WHOLLY_BLIND because of historical exposure. Engineering acceptance does not replace the historical official CVDP result of 243/302 (Opus 4.8 · plugin v1.2.63).

Acceptance checkpoint · plugin v1.17.71

The canonical engineering-acceptance checkpoint was recorded on v1.17.71. Historical AI author records are mixed; this checkpoint is not attributed to one guessed model alias.

Latest pending-case replay · plugin v1.17.75

The latest read-only replay covered only the single pending case on plugin v1.17.75 with vibeic-eda 0.3.46. It did not re-run or re-accept the full subset.

1 unresolved case

cvdp_copilot_64b66b_decoder_0011 remains unresolved: the public specification lacks a recognized-type payload-error predicate. It remains in the denominator.

This bounded progress bundle publishes normalized acceptance and four-stage records, selected supporting receipts, and source/published SHA-256 hashes. Full RTL, the executable challenge corpus, the dataset and unavailable raw transcripts are not included; this is not a complete reproducible benchmark package.

Validated end-to-end on 8 real ICs

Across 6 design classes — cryptographic primitive, secure processor, RISCV CPU (Verilog / SystemVerilog / VHDL / SpinalHDL), RISCV SoC, mixed-signal ΔΣ-ADC, and an edge-AI inference accelerator. Every step runs real open-source signoff inside the vibeic-eda container — yosys + OpenROAD + klayout + magic + netgen + ngspice + iverilog+SDF + SymbiYosys.

2 of 12 (IC × PDK) cells converged

IC PDK State Evidence
spmIHP-SG13G2in progressno landed evidence yet
spmsky130AconvergedPASS_WITH_WAIVERS · completion audit
spmGF180MCUconvergedPASS_WITH_WAIVERS · completion audit
sha256sky130Ain progressno landed evidence yet
caravel_user_projectsky130Ain progressno landed evidence yet
edge_llm_accelNanGate45in progressno landed evidence yet
edge_llm_matmul_accelNanGate45in progressno landed evidence yet
ibexsky130Ain progressno landed evidence yet
opentitan_aessky130Ain progressno landed evidence yet
subservientsky130Ain progressno landed evidence yet
subservientGF180MCUin progressno landed evidence yet
u_hawaii_adcsky130Ain progressno landed evidence yet

A cell counts as converged only when its landed evidence says so — the verdict is read from the run's own completion audit, not from a status table. Cells with no landed evidence show as in progress; they are never counted as passing and never removed from the denominator. Updated 2026-10-09 06:17 UTC+08:00.

Crypto / Secure-proc

SHA-256 ×3 variants · SPM

RISCV CPU

PicoRV32 · CV32E40P · Ibex · SERV

RISCV SoC

Subservient (SERV-SoC) · NEORV32 (VHDL via GHDL)

Mixed-signal

U-Hawaii ΔΣ-ADC — A1-A9 + M1-M4 tracks exercised on real ngspice/klayout artefacts

Documented edge cases

DarkRISCV (BRAM-as-flops over-utilisation) · VexRiscv (SpinalHDL out of scope)

First formal-verification PASS

SERV via SymbiYosys + RVFI BMC depth=10 on the real RTL

Edge-AI accelerator

edge_llm_accel — 1.36M-cell INT4 GEMM engine, the 8th benchmark IC (below)

8th benchmark IC · first run on a NEW PDK — NanGate45 / FreePDK45 (open 45nm)

edge_llm_accel — a 1.36M-cell edge-LLM accelerator in ~14 hours

Prompted by the Kimi K3 “48-hour chip” demo, we designed our OWN same-class chip — a 64×64 weight-stationary INT4 systolic GEMM core (4096 MAC/cycle) with a 20-bank ~195 KB SRAM scratchpad and fused dequantization — and drove it docs→GDS through the standard Vibe-IC front door on the SAME open PDK (NanGate45), the first time this PDK ran in our flow. 3/3 RTL modules generated from the L1–L9 design documents, zero reused IP. Doc-driven dual-track verification caught a real pre-silicon RTL bug (a half-rate weight-load chain that silently degraded the array to 32×64 and leaked state across runs) — fixed and re-proven bit-true before synthesis.

Metric Kimi K3 demo Vibe-IC edge_llm_accel
Std cells1.46M1,356,030 + 20 SRAM macros (93%)
Design area3.981 mm²3.07 mm² @ 55% util (die 5.76 mm²)
Clock100 MHz100 MHz MET · SPEF WNS +1.08 ns · TNS 0
Route DRC / antenna—0 / 0
Functional verification—180 random 64×64 tiles bit-true · 9 seeds · 1 real bug caught pre-silicon
Wall-clock (docs → GDS, incl. all convergence)48 h~14 h

Honest scope: NanGate45 / FreePDK45 is a simulation-grade, non-foundry 45nm enablement (fictional process, no LVS deck, abstract FakeRAM SRAM macros) — so this is synth → macro place → PnR → CTS → detailed-route-DRC-clean → GDS, i.e. “tape-out simulation”, the SAME level as the Kimi demo; real foundry sign-off is demonstrated separately on sky130A / GF180MCU / IHP-SG13G2 / the commercial 180nm PDK above. The educational FreePDK45 KLayout deck reports 23,082 items, every one attributed (15,814 = the deck’s 200 nm well-separation reading vs the library’s abutting-row wells; 7,247 = its simplified flat antenna model where OpenROAD’s hierarchical check reports 0; 20 = exactly the 20 abstract SRAM macros; 1 metal item). Autonomy models differ: Kimi = one LLM iterating alone for 48 h; Vibe-IC = a deterministic runner chain + gated AI convergence — and the run distilled 11 chip-agnostic tool/flow fixes back into the plugin. Evidence note: this historical result has not yet landed in the separate public benchmark-data repository. Every figure is measured, not claimed.

Commercial-PDK sign-off · a 180nm NDA foundry PDK · native, no Calibre license

First 4 ICs on a commercial PDK

Beyond the open-PDK runs, the same ICs are driven through a real commercial 180nm foundry flow. DRC runs the foundry's own Calibre .rule deck (224 layers / 4533 rules) NATIVELY on the forked KLayout engine (svrfdrc) — no Calibre license. Honest status: two digital ICs reach production-grade sign-off; the analog ΔΣ-ADC reaches real corner-simulation; the crypto core's backend is clean and its only residuals are documented open-source-tool scale floors — never silently dropped.

IC Class GDSII Sign-off DRC LVS STA Status
spm SPI peripheral 2.16 MB 4533 / 4533 MATCH MET · +5.55 ns Converged
subservient RISC-V SoC (SERV) 6.3 MB 4533 / 4533 MATCH MET · +2.73 ns Converged
u_hawaii_adc ΔΣ-ADC + LDO (analog) ldo 432,726 B · ΔΣ 79,410 B A6 downstream — 9-corner PVT A8 hardmacro GDS
sha256 Crypto hash 84–90 MB 4528 / 4533 dev+net match MET · +17.06 ns Backend green

The sign-off target is a 180nm commercial NDA foundry PDK. spm & subservient: PASS_WITH_WAIVERS — GDSII + native DRC 4533/4533 + full LVS MATCH (KLayout NetlistComparer + netgen, 0 power shorts) + STA MET; only an FPGA-board hardware waiver remains. u_hawaii_adc: the LDO reaches a full 9-corner TT/SS/FF × (−40/27/125 °C) real-ngspice PVT sweep (Vout ≈ 1.199 V) on BOTH open IHP-SG13G2 and the commercial 180nm PDK; it walls downstream at A6 layout parasitic-verification (real layout pending). sha256: backend is clean (0 router-DRC, STA MET) — the 5 firing sign-off-DRC rules are 4 FEOL over-fires on foundry-qualified std-cell interiors plus 1 metal-density gap; LVS matches device-for-device (108,150) and net-for-net (55,541) with 3 top-level pins a known NetlistComparer artifact; and the single-threaded svrfdrc runtime (~6 h on 90 MB) is a documented open-source-tool scale floor — a spatial-tiling parallelization is in progress. Every figure is measured, not claimed.

Open-PDK Benchmark IC campaign · sky130A / GF180MCU / IHP-SG13G2 / NanGate45 · 2026-07

An IC × PDK convergence matrix, updated as it converges

A wider matrix than the single-PDK results above: the same 9 designs driven across 4 open PDKs. Each (IC × PDK) cell is a distinct, independently-graded result. Status here is re-derived from raw run artifacts, never from a run's own self-report.

IC × PDK GDS Sign-off DRC LVS STA Status
spm × IHP-SG13G2 822,084 B 592 / 592 MATCH MET · +6.14 / +0.17 ns Converged
spm × sky130A 730,458 B 145 / 145 MATCH MET · +4.56 / +0.33 ns Converged
spm × GF180MCU 1,180,456 B 763 / 763 MATCH MET · +1.73 / +0.57 ns Converged
sha256 × sky130A present—— setup gap Converging
caravel_user_project × sky130A ———— Converging
edge_llm_accel × NanGate45 ———— Converging
edge_llm_matmul_accel × NanGate45 ———— Converging
ibex × sky130A ———— Converging
opentitan_aes × sky130A present—— setup gap Converging
subservient × sky130A ———— Converging
subservient × GF180MCU ———— Converging
u_hawaii_adc × sky130A analog track —— PVT sweep Converging

Progress 2026-07-31 — see text.

Historical snapshot: 3 of 12 cells had been independently re-derived PASS as of 2026-07-24: spm × IHP-SG13G2 (DRC 592/592) and spm × sky130A (DRC 145/145) — the same "checked-clean/total" claim strength as the 4533/4533 commercial-PDK number, just a smaller deck (a property of open vs. commercial PDKs, not a thinner check). At that snapshot, the other 9 cells remained in active convergence with genuine, documented residuals. Current evidence is linked in the live matrix above and in the public vibeic/benchmark-data repository.

Why Vibe-IC?

Lower the barrier from decades of training to a conversation.

Traditional IC Design

  • 10+ years experience required
  • $1M+ commercial EDA tools
  • NDA-locked PDK
  • 6–12 months design cycle
  • $100K+ per tapeout
vs

Vibe-IC

  • Anyone — natural language
  • Open-source EDA — 48 EDA tools in one Docker
  • Open + Custom PDK — SKY130, GF180, IHP-SG13G2 + NanGate45 / ASAP7 (research-only) + any commercial PDK
  • 2–3 months to tapeout-ready
  • $10K (Efabless chipIgnite)

That column compares Vibe-IC to the traditional commercial flow. For a comparison against the two open projects it gets named next to most often — OpenECOS and OpenROAD — including where Vibe-IC is measurably behind them, see Similar projects.