Evaluation · official scores, engineering progress + IC sign-off
Official benchmark scores use clean-room blind runs and each benchmark's official scorer. Engineering checkpoints are reported separately with their scope, exposure history and unresolved cases. The IC results document end-to-end spec → GDSII work through open-source sign-off; each result retains its evidence and limitations.
Open Benchmarks · fully blind · official upstream testbenches
Every open-benchmark score below follows clean-room blind rules: the authoring agent sees only the permitted prompt/context — never the hidden testbench, golden, or sibling solutions — and the benchmark's official scorer establishes the result. Each cell names the AI final semantic authority and Vibe-IC plugin version. The newest completed results are the 2026-09-06 single-shot blind runs on shipped v1.17.60 (plugin tree 1eb2a241) with Claude Opus 5 as final reviewer for VerilogEval-v2 (152/156) and VerilogEval-Human (154/156, tree a938bc0a), and Claude Fable 5 on shipped v1.15.20 for RTLLM (48/50, 2026-09-01). Beside each VerilogEval score we state the theoretical maximum: problems proven broken with re-runnable evidence — a golden that fails its own testbench, or contradicts its own prompt — stay in the denominator and are named, never rewritten. All evidence cells are public.
Measured datapoints · model + plugin identified
All rows use the canonical clean-room entry and official scorer, but they were not run on one identical plugin build. Read each cell as an AI-authority + plugin datapoint, not as a controlled model-only A/B. Single-shot pass@1 and terminal loop2converge are never merged. The 2026-08-31 VE cells started from v1.13.78 plus four disclosed campaign paths, so they are not shipped-release reproductions; their complete public evidence is linked below.
| Benchmark · blind · official scorer | Claude | GPT | Kimi |
|---|---|---|---|
| VerilogEval-v2 (156) | 152/156 pass@1 = max 152 Claude Opus 5 · shipped v1.17.60 · 2026-09-06 |
151/156 pass@1 → 153/156 terminal GPT-5.6-Sol · campaign-patched v1.13.78 |
147/156 single Kimi K3 · plugin version not bound |
| VerilogEval-Human (156) | 154/156 pass@1 = max 154 Claude Opus 5 · tree a938bc0a · 2026-09-06 |
152/156 pass@1 → 153/156 terminal GPT-5.6-Sol · campaign-patched v1.13.78 |
149/156 single Kimi K3 · plugin version not bound |
| RTLLM v2.0 (50) | 48/50 pass@1 Claude Fable 5 · v1.15.20 · 46/48 discriminating |
48/50 pass@1 GPT-5.6-Sol · v1.14.5 · 45/47 discriminating |
Not measured |
Read-out: the 2026-09-06 cells on shipped v1.17.60 are single-shot blind runs. VerilogEval-v2 records 152/156, which is its theoretical maximum: four problems are proven broken with re-runnable evidence in each cell's theoretical_max.json — Prob099 (the golden cannot compile against its own testbench), Prob062, Prob093 and Prob149 (the golden contradicts what the prompt states). VerilogEval-Human records 154/156, its theoretical maximum (Prob062 and Prob093). On RTLLM the newest published cell (Claude Fable 5 · shipped v1.15.20, 2026-09-01) records 48/50 official pass@1 (46/48 discriminating), matching GPT-5.6-Sol · v1.14.5's 48/50; its pre-scorer Program First + AI review layer accepted 50/50, and that semantic layer does not rewrite the official score. Its two residuals are one proven dataset defect (the radix2_div golden fails its own testbench, 3/8) and one plugin-gate misdirection captured as public issue vibe-ic#1998. The earlier 153/156 on VerilogEval-v2 included Prob149, which was passable only through a lessons entry later found to be oracle-derived and withdrawn. Denominators are always the original ones.
Plus: CVDP 243/302 = 80.46% official-compliant blind pass@1 on the public no-commercial code-generation set (measured on Opus 4.8 · plugin v1.2.63 — a separate earlier campaign) — prompt+context-only, with the hidden cocotb harness and golden kept off-limits; and a MetRex premise-dissolving demo — instead of predicting post-synthesis area, Vibe-IC runs MetRex's exact Yosys + sky130 recipe and reproduces its golden area to a 1.15% median (90% within 5%). Honest caveats included: dataset, specification, and scorer/tool-dialect limitations are recorded with the exact official score, never silently dropped.
Engineering progress · 2026-09-06 · original unfinished69 subset
Engineering acceptance is 68/69 on the original unfinished69 hard-tail subset, with 1 case unresolved. This checkpoint covers that fixed subset only. It is not a full-dataset score, Pass@1, or an all-pass result.
Official score: NOT_RUN. Blindness status: NOT_ELIGIBLE_AS_WHOLLY_BLIND because of historical exposure. Engineering acceptance does not replace the historical official CVDP result of 243/302 (Opus 4.8 · plugin v1.2.63).
The canonical engineering-acceptance checkpoint was recorded on v1.17.71. Historical AI author records are mixed; this checkpoint is not attributed to one guessed model alias.
The latest read-only replay covered only the single pending case on plugin v1.17.75 with vibeic-eda 0.3.46. It did not re-run or re-accept the full subset.
cvdp_copilot_64b66b_decoder_0011 remains unresolved: the public specification lacks a recognized-type payload-error predicate. It remains in the denominator.
This bounded progress bundle publishes normalized acceptance and four-stage records, selected supporting receipts, and source/published SHA-256 hashes. Full RTL, the executable challenge corpus, the dataset and unavailable raw transcripts are not included; this is not a complete reproducible benchmark package.
Across 6 design classes — cryptographic primitive, secure processor, RISCV CPU (Verilog / SystemVerilog / VHDL / SpinalHDL), RISCV SoC, mixed-signal ΔΣ-ADC, and an edge-AI inference accelerator. Every step runs real open-source signoff inside the vibeic-eda container — yosys + OpenROAD + klayout + magic + netgen + ngspice + iverilog+SDF + SymbiYosys.
| IC | PDK | State | Evidence |
|---|---|---|---|
| spm | IHP-SG13G2 | in progress | no landed evidence yet |
| spm | sky130A | converged | PASS_WITH_WAIVERS · completion audit |
| spm | GF180MCU | converged | PASS_WITH_WAIVERS · completion audit |
| sha256 | sky130A | in progress | no landed evidence yet |
| caravel_user_project | sky130A | in progress | no landed evidence yet |
| edge_llm_accel | NanGate45 | in progress | no landed evidence yet |
| edge_llm_matmul_accel | NanGate45 | in progress | no landed evidence yet |
| ibex | sky130A | in progress | no landed evidence yet |
| opentitan_aes | sky130A | in progress | no landed evidence yet |
| subservient | sky130A | in progress | no landed evidence yet |
| subservient | GF180MCU | in progress | no landed evidence yet |
| u_hawaii_adc | sky130A | in progress | no landed evidence yet |
A cell counts as converged only when its landed evidence says so — the verdict is read from the run's own completion audit, not from a status table. Cells with no landed evidence show as in progress; they are never counted as passing and never removed from the denominator. Updated 2026-10-09 06:17 UTC+08:00.
SHA-256 ×3 variants · SPM
PicoRV32 · CV32E40P · Ibex · SERV
Subservient (SERV-SoC) · NEORV32 (VHDL via GHDL)
U-Hawaii ΔΣ-ADC — A1-A9 + M1-M4 tracks exercised on real ngspice/klayout artefacts
DarkRISCV (BRAM-as-flops over-utilisation) · VexRiscv (SpinalHDL out of scope)
SERV via SymbiYosys + RVFI BMC depth=10 on the real RTL
edge_llm_accel — 1.36M-cell INT4 GEMM engine, the 8th benchmark IC (below)
8th benchmark IC · first run on a NEW PDK — NanGate45 / FreePDK45 (open 45nm)
Prompted by the Kimi K3 “48-hour chip” demo, we designed our OWN same-class chip — a 64×64 weight-stationary INT4 systolic GEMM core (4096 MAC/cycle) with a 20-bank ~195 KB SRAM scratchpad and fused dequantization — and drove it docs→GDS through the standard Vibe-IC front door on the SAME open PDK (NanGate45), the first time this PDK ran in our flow. 3/3 RTL modules generated from the L1–L9 design documents, zero reused IP. Doc-driven dual-track verification caught a real pre-silicon RTL bug (a half-rate weight-load chain that silently degraded the array to 32×64 and leaked state across runs) — fixed and re-proven bit-true before synthesis.
| Metric | Kimi K3 demo | Vibe-IC edge_llm_accel |
|---|---|---|
| Std cells | 1.46M | 1,356,030 + 20 SRAM macros (93%) |
| Design area | 3.981 mm² | 3.07 mm² @ 55% util (die 5.76 mm²) |
| Clock | 100 MHz | 100 MHz MET · SPEF WNS +1.08 ns · TNS 0 |
| Route DRC / antenna | — | 0 / 0 |
| Functional verification | — | 180 random 64×64 tiles bit-true · 9 seeds · 1 real bug caught pre-silicon |
| Wall-clock (docs → GDS, incl. all convergence) | 48 h | ~14 h |
Honest scope: NanGate45 / FreePDK45 is a simulation-grade, non-foundry 45nm enablement (fictional process, no LVS deck, abstract FakeRAM SRAM macros) — so this is synth → macro place → PnR → CTS → detailed-route-DRC-clean → GDS, i.e. “tape-out simulation”, the SAME level as the Kimi demo; real foundry sign-off is demonstrated separately on sky130A / GF180MCU / IHP-SG13G2 / the commercial 180nm PDK above. The educational FreePDK45 KLayout deck reports 23,082 items, every one attributed (15,814 = the deck’s 200 nm well-separation reading vs the library’s abutting-row wells; 7,247 = its simplified flat antenna model where OpenROAD’s hierarchical check reports 0; 20 = exactly the 20 abstract SRAM macros; 1 metal item). Autonomy models differ: Kimi = one LLM iterating alone for 48 h; Vibe-IC = a deterministic runner chain + gated AI convergence — and the run distilled 11 chip-agnostic tool/flow fixes back into the plugin. Evidence note: this historical result has not yet landed in the separate public benchmark-data repository. Every figure is measured, not claimed.
Commercial-PDK sign-off · a 180nm NDA foundry PDK · native, no Calibre license
Beyond the open-PDK runs, the same ICs are driven through a real commercial 180nm foundry flow. DRC runs the foundry's own Calibre .rule deck (224 layers / 4533 rules) NATIVELY on the forked KLayout engine (svrfdrc) — no Calibre license. Honest status: two digital ICs reach production-grade sign-off; the analog ΔΣ-ADC reaches real corner-simulation; the crypto core's backend is clean and its only residuals are documented open-source-tool scale floors — never silently dropped.
| IC | Class | GDSII | Sign-off DRC | LVS | STA | Status |
|---|---|---|---|---|---|---|
| spm | SPI peripheral | 2.16 MB | 4533 / 4533 | MATCH | MET · +5.55 ns | Converged |
| subservient | RISC-V SoC (SERV) | 6.3 MB | 4533 / 4533 | MATCH | MET · +2.73 ns | Converged |
| u_hawaii_adc | ΔΣ-ADC + LDO (analog) | ldo 432,726 B · ΔΣ 79,410 B | A6 downstream | — | 9-corner PVT | A8 hardmacro GDS |
| sha256 | Crypto hash | 84–90 MB | 4528 / 4533 | dev+net match | MET · +17.06 ns | Backend green |
The sign-off target is a 180nm commercial NDA foundry PDK. spm & subservient: PASS_WITH_WAIVERS — GDSII + native DRC 4533/4533 + full LVS MATCH (KLayout NetlistComparer + netgen, 0 power shorts) + STA MET; only an FPGA-board hardware waiver remains. u_hawaii_adc: the LDO reaches a full 9-corner TT/SS/FF × (−40/27/125 °C) real-ngspice PVT sweep (Vout ≈ 1.199 V) on BOTH open IHP-SG13G2 and the commercial 180nm PDK; it walls downstream at A6 layout parasitic-verification (real layout pending). sha256: backend is clean (0 router-DRC, STA MET) — the 5 firing sign-off-DRC rules are 4 FEOL over-fires on foundry-qualified std-cell interiors plus 1 metal-density gap; LVS matches device-for-device (108,150) and net-for-net (55,541) with 3 top-level pins a known NetlistComparer artifact; and the single-threaded svrfdrc runtime (~6 h on 90 MB) is a documented open-source-tool scale floor — a spatial-tiling parallelization is in progress. Every figure is measured, not claimed.
Open-PDK Benchmark IC campaign · sky130A / GF180MCU / IHP-SG13G2 / NanGate45 · 2026-07
A wider matrix than the single-PDK results above: the same 9 designs driven across 4 open PDKs. Each (IC × PDK) cell is a distinct, independently-graded result. Status here is re-derived from raw run artifacts, never from a run's own self-report.
| IC × PDK | GDS | Sign-off DRC | LVS | STA | Status |
|---|---|---|---|---|---|
| spm × IHP-SG13G2 | 822,084 B | 592 / 592 | MATCH | MET · +6.14 / +0.17 ns | Converged |
| spm × sky130A | 730,458 B | 145 / 145 | MATCH | MET · +4.56 / +0.33 ns | Converged |
| spm × GF180MCU | 1,180,456 B | 763 / 763 | MATCH | MET · +1.73 / +0.57 ns | Converged |
| sha256 × sky130A | present | — | — | setup gap | Converging |
| caravel_user_project × sky130A | — | — | — | — | Converging |
| edge_llm_accel × NanGate45 | — | — | — | — | Converging |
| edge_llm_matmul_accel × NanGate45 | — | — | — | — | Converging |
| ibex × sky130A | — | — | — | — | Converging |
| opentitan_aes × sky130A | present | — | — | setup gap | Converging |
| subservient × sky130A | — | — | — | — | Converging |
| subservient × GF180MCU | — | — | — | — | Converging |
| u_hawaii_adc × sky130A | analog track | — | — | PVT sweep | Converging |
Progress 2026-07-31 — see text.
Historical snapshot: 3 of 12 cells had been independently re-derived PASS as of 2026-07-24: spm × IHP-SG13G2 (DRC 592/592) and spm × sky130A (DRC 145/145) — the same "checked-clean/total" claim strength as the 4533/4533 commercial-PDK number, just a smaller deck (a property of open vs. commercial PDKs, not a thinner check). At that snapshot, the other 9 cells remained in active convergence with genuine, documented residuals. Current evidence is linked in the live matrix above and in the public vibeic/benchmark-data repository.
Lower the barrier from decades of training to a conversation.
That column compares Vibe-IC to the traditional commercial flow. For a comparison against the two open projects it gets named next to most often — OpenECOS and OpenROAD — including where Vibe-IC is measurably behind them, see Similar projects.