Implementation review · v1.12.34 · 2026-08-28
v1.12.34 has real PPA records, parsers, feasibility checks, search contracts and seven PPA skills. Three canonical edges remeasure after an action and two of them are now rollback-proven; no edge is controller-bound yet. A feasibility audit says why: of the 21 declared edges, the two that name a metric cannot reach the producer that would have to change, and 19 name no metric at all. This page separates those facts from the plan.
This is the 2026-08-20 snapshot that motivated the five-phase plan, not the current implementation status. At v1.11.7, parsing flow/phase1_phase2_phase3.yaml found 19 closed_loop declarations. The v1.12.34 review below supersedes every present-tense conclusion from this baseline.
At v1.11.7: steps 10→7, 20→19, 23→32 and 32→32 were the four performance-related declarations. Their execution-proof levels had not yet been audited.
At v1.11.7, only IR-drop 24→15 was classified on the power axis; step 33 recorded total power without a declared fallback edge.
At v1.11.7, the audit found zero area-axis fallback declarations and no declared area metric. Later versions changed this; see the v1.12.34 review.
At that baseline, step 32 declared an aggregator fallback over steps 21–28 after ECO. The declaration was the reason to investigate the lane, not proof that every downstream step was rerun or that rollback worked; those proof levels are measured separately below.
| Program | Flow references | Status |
|---|---|---|
| ppa_head_to_head_check.py | 0 | Gate exists in the hygiene suite, marked run_tolerating_uncheckable, and its fixture pair is recorded as NOT_WRITTEN_YET debt. |
| ppa_predict_aggregate.py | 0 | Pre-synthesis power / area / Fmax estimator. Unreachable from the flow. |
| ppa_area_threshold_check.py | 0 | Synthesises an original / optimised pair with the same recipe and compares cells and wires — exactly the gate the area axis needs. |
| readme_ppa_extractor.py | 0 | Reads the PPA targets the design declares in its own documents. |
In that v1.11.7 snapshot, the ppa-predict skill was also referenced zero times. v1.12.34 now gives it a mandatory deterministic aggregator; that is documented in the current review.
Read out of the sources, not from memory: OpenROAD a47e38e, LibreLane bf8cc13, OpenROAD-flow-scripts 3476215. Their PPA is three layers — the actuators inside the tool, the knobs the flow exposes, and a search that drives them.
The tool already trades one axis against another, and we use almost none of it.
repair_timing [-setup] [-hold] [-recover_power percent_of_paths_with_slack]
[-sequence move_list] [-skip_pin_swap] [-skip_gate_cloning]
# -recover_power spends leftover timing slack on power. That is P↔P, in the tool.
set_opt_config [-limit_sizing_area] [-limit_sizing_leakage]
[-keep_sizing_site] [-keep_sizing_vt]
[-sizing_area_limit] [-sizing_leakage_limit]
# a ceiling on what fixing timing is allowed to cost in area and in leakage.
replace_arith_modules [-path_count n] [-slack_threshold f] [-target opto_goal]
# swaps arithmetic implementations on critical paths. Micro-architecture, in the tool.
Its metrics are finer-grained than ours too: power splits into internal / switching / leakage / total, and area into stdcell / macros / padcells / cover.
Two families matter most, because they are the timing-versus-area trade made adjustable instead of hard-coded.
PL_RESIZER_SETUP_MAX_UTIL_PCT core area fixing setup is allowed to consume PL_RESIZER_HOLD_MAX_UTIL_PCT core area fixing hold is allowed to consume PL_RESIZER_SETUP_REPAIR_TNS_PCT what fraction of violating endpoints to even try # and the same fourteen again, prefixed GRT_, for the post-global-route pass.
It also ships two exploration flows — but read what they actually are:
design__instance__area, then tries FP_CORE_UTIL=99 and falls back to 40 on failure. Its own header calls it “a custom demo flow”, and it optimises a single objective.This is the part worth stealing. Ray Tune drives it, and the objective is explicitly a head-to-head against a reference run.
coeff_perform, coeff_power, coeff_area = 10000, 100, 100 eff_clk_period = clk_period - worst_slack # only when slack is negative performance = percent(eff_clk_period_ref, eff_clk_period) power = percent(reference["total_power"], metrics["total_power"]) area = percent(100 - reference["final_util"], 100 - metrics["final_util"]) ppa = performance*10000 + power*100 + area*100 score = (upper_bound - ppa) * (step/100)**-1 + (ppa/10) * num_drc
Four decisions in there are worth naming:
step is a divisor, so a run that dies early or leaves violations cannot buy a good score._SDC_CLK_PERIOD).The space it searches, per design:
_SDC_CLK_PERIOD float [3.5, 7.0] CORE_UTILIZATION int [20, 50] CORE_ASPECT_RATIO float [0.5, 2.0] CELL_PAD_IN_SITES_GLOBAL_PLACEMENT int [0, 3] CELL_PAD_IN_SITES_DETAIL_PLACEMENT int [0, 3] PLACE_DENSITY_LB_ADDON float [0.0, 0.2] CTS_CLUSTER_SIZE int [10, 200] CTS_CLUSTER_DIAMETER int [20, 400] _FR_LAYER_ADJUST float [0.1, 0.3]
In the pinned OpenROAD, LibreLane and ORFS source revisions listed below, the exposed tuner knobs are place-and-route knobs. AutoTuner does not touch synthesis strategy, RTL or micro-architecture; LibreLane's synthesis exploration uses nine discrete presets and does not feed placement. None of those three pinned flows performs a unified cross-layer search.
Five phases. Each one is useful on its own and each one is a precondition for the next, which is why they are numbered rather than listed. Every phase closes on a command that returns zero only when the work is really on main.
PowerArea
Nothing downstream can converge on a metric the flow does not record. Add the metrics as declared required_outputs on the steps that already produce them, using OpenROAD's own names so a reader can cross-check us against ORFS without translation.
step 9 design__instance__area, __count
step 17 design__instance__utilization
design__core__area
step 21 route__wirelength
step 33 power__internal|switching
|leakage|total (split)
step 37 design__die__area
Every unmeasured field is the literal NOT_MEASURED. A plausible default here would hide precisely the hole this phase exists to expose.
PowerArea
This is the phase the whole plan is named after. Two new closed_loop declarations, in the same shape the timing edges already use — plus a ceiling on what fixing timing may cost, which is the trade LibreLane exposes and we do not.
step 33 closed_loop: fallback_to: 17
trigger: power__total above the
design's declared budget
step 9 gate: ppa_area_threshold_check
trigger: area above the ceiling
step 32 + RESIZER_SETUP_MAX_UTIL_PCT
+ RESIZER_HOLD_MAX_UTIL_PCT
+ repair_timing -recover_power
ppa_area_threshold_check.py already synthesises an original / optimised pair and compares cells and wires. It needs wiring, not writing.
PerformancePowerArea
ppa_head_to_head_check.py opens by stating the problem better than we can: our published numbers prove an AI can produce relatively correct RTL; they do not prove it can produce better silicon, and a reviewer is entitled to answer them with “so what?”
arm A vibe-ic our flow, our defaults arm B LibreLane Classic, its defaults arm C ORFS + AutoTuner, tuned best # arm C is the one that makes the # comparison worth reading.
Beating an untuned baseline proves nothing. The claim only means something against a flow that was allowed to search. What is missing is the fixture pair, a flow step that calls the program, and the third arm.
PerformancePowerArea
AutoTuner's bottleneck is that a good search wants hundreds of full flow runs; they solve it with a Ray cluster. We already have a fleet and a dispatch discipline. That is a structural fit, not a coincidence.
Port the objective, with one deliberate change: the weights become a declared input, not a constant.
# ORFS hard-codes 10000/100/100. A 100:1 preference for speed over power is a # value judgement about a market, and it belongs to the design, not to us. ppa_weights: performance: <from L19, no default> power: <from L19, no default> area: <from L19, no default> # A design that declares no preference inherits ORFS's ratio EXPLICITLY, recorded # in the report as "inherited, not chosen", so a reader can see whose call it was.
Keep the two anti-cheating terms verbatim, because they are the reason the score cannot be gamed: num_drc as a penalty, and (step/100)**-1 so a run that stops early is scored on how far it got.
PerformancePowerArea
Everything in Phase 3 is a place-and-route knob, because that is all the open flows tune. The search space we can offer that they cannot:
| Layer | Lever | Who searches it today |
|---|---|---|
| Place & route | utilisation, density, padding, CTS clustering | ORFS AutoTuner |
| Synthesis | SYNTH_STRATEGY | LibreLane — 9 fixed presets, a human picks |
| Arithmetic | adder / multiplier architecture | OpenROAD, narrowly — replace_arith_modules |
| RTL | pipelining, resource sharing, encoding | none in the pinned survey |
| Micro-architecture | the structure the spec never pinned | none in the pinned survey |
The bottom two rows motivate the plan. An agent can propose RTL and micro-architecture changes, but an accepted rewrite must be required to pass equivalence, lint and STA and to preserve their verdicts. v1.12.34 does not yet prove that this gate set executes for every cross-layer candidate.
A PPA number is a claim about silicon. These are the ways such a claim goes wrong, and every one of them has precedent in this repository.
NOT_MEASURED and blocks. “We did not look” and “we looked and it was fine” must not produce the same artefact.0.5ic/d3, which is red for an unrelated reason: step 0.5ic is newer than every published cell, so no published run has produced its declared outputs.Sources read for this plan: flow/phase1_phase2_phase3.yaml @ v1.11.7 · OpenROAD a47e38e · librelane bf8cc13 · OpenROAD-flow-scripts 3476215. Every count in the first section was produced by parsing the flow, not recalled.
The page above is the v1.11.7 plan written on 2026-08-20. This appended review supersedes its present-tense status claims; it does not rewrite that historical plan.
v1.12.34 keeps the canonical PPA records, scoped identities, provenance, hard feasibility, Pareto comparison, search budgets, closure state machines, five report parsers and seven PPA skills. It adds a more important property around them: every top-level program is now reachable, all ten checker-shaped PPA programs have machine runners, and the generated PPA report has a claims file plus a gate that refuses prose stronger than its evidence.
Reachable is not the same as bound. The dedicated registry still reports 21 declared edges, zero BOUND controllers and 21 DECLARED_ONLY edges; its one executable controller is bound to no canonical edge. The generic search still records externally produced trials instead of launching a complete cross-layer campaign itself.
REMEASURED means that an action ran and the relevant condition was measured again. It does not by itself mean converged, accepted, shipped or rollback-proven. The tiers are nested, so an edge counted at a higher tier is no longer counted at a lower one: 4→1 (functional simulation) is the one REMEASURED edge, and the two PPA-related ones, 23→32 and 32→32, have since been promoted to ROLLBACK_PROVEN.
The release closes reproducibility and wiring gaps around PPA. It does not promote declaration into convergence.
| Area | v1.12.34 contract | Honest boundary |
|---|---|---|
| Program reachability | The wiring audit covers 1,301 top-level programs and 657 checker-shaped ones; all ten ppa_* checkers resolve to a machine runner — eight from the hygiene lane, one (head-to-head) also from the flow yaml, and two from another program. | A separate automatic-verdict audit still lists 27 non-PPA gates outside an automatic verdict path. Reachability is not proof that every program is a canonical flow gate. |
| ppa-diagnose | The skill now runs ppa_diagnostic_router.py first, continues only on an explicit HANDOFF, then builds a path-and-hash evidence context with ppa_agent_context_build.py. | PROGRAM_DECIDED, REFUSED and UNDETERMINED stop the AI path. Missing evidence may not be replaced by a fluent hypothesis. |
| ppa-measure | The skill now requires ppa_signoff_records and ppa_eco_spare_records to produce canonical records; ppa_report_gen then produces report.md plus claims.json, and ppa_page_claim_check checks citations, numbers and evidence strength. | The two record producers are operator-run today, not canonical flow clauses: spare_preservation.json is conditional and undeclared, so making it an unconditional dependency would falsely red every design without spares. |
| ppa-predict | The skill now makes ppa_predict_aggregate.py own the cell-count-to-area/power/Fmax arithmetic, PDK constants, ranges, status labels and assumptions. | A missing or zero cell-count anchor is rc 2 CANNOT CHECK. The result stays ESTIMATED / RTL_PROXY and is explicitly ineligible for physical-PPA claims. |
| Published-report claims | The repository's winner report passes over 35 sentences, 139 claims and 9 banned unqualified forms. PPA mutation fixtures prove the good arm passes and a promoted or broken claim fails. | That gate protects the generated repository report/claims pair, not this static marketing page. This page is instead pinned to the source commit and verification receipt below. |
The vocabulary reaches beyond place and route, but declaration, execution and proof are separate levels.
| Layer | Current lever surface | Relevant flow edges | Measured status |
|---|---|---|---|
| RTL | pipelining and state encoding; specification permission and equivalence mode are recorded | 2→1 · 3→1 · 4→1 · 5→1 | 4→1 REMEASURED; the others DECLARED_ONLY |
| Synthesis and constraints | AREA 0..3 and DELAY 0..4 strategy vocabulary | 8→7 · 9→1 · 10→7 · 13→9 · 14→9 | DECLARED_ONLY |
| Arithmetic | adder and multiplier architecture choices are declared; Vibe-IC has no candidate rewrite executor yet | cross-layer search space, not a bound canonical edge | DECLARED_ONLY |
| Microarchitecture | module hierarchy only; not general microarchitecture exploration | cross-layer search space, not a bound canonical edge | DECLARED_ONLY |
| Analog | post-extraction and hardware-correlation fallback declarations | A7→A3 · A9→A3 | DECLARED_ONLY |
| Physical design and sign-off | floorplan, placement, ECO spares, CTS, routing and constraint vocabulary | 20→19 · 23→32 · 24→15 · 25/26/27→21 · 28→15 · 31→32 · 32→32 · 33→17 | 23→32 and 32→32 ROLLBACK_PROVEN — the undo has a named test |
| System | What it already does | Vibe-IC's position |
|---|---|---|
| OpenROAD | The real implementation actuator: timing repair, buffering, sizing, VT swaps and narrowly scoped arithmetic replacement. Arithmetic replacement currently implements setup optimization; hold, power and area targets remain unavailable. | Vibe-IC does not replace OpenROAD's optimizer; it adds evidence identity, provenance, feasibility and controlled adoption around it. |
| LibreLane | SynthesisExploration runs nine synthesis presets and reports them for a human choice. Its separate Optimizing demo selects the smallest-area result from three presets, then tries high utilisation with a lower-utilisation fallback. | Vibe-IC has stricter scope and evidence contracts, but has not yet connected them to an equally complete trial executor. |
| ORFS AutoTuner | Actually launches and manages EDA trials through Ray Tune. The current CLI exposes hyperopt, ax, optuna, pbt and random, with reference-relative PPA terms plus early-stop and DRC penalties. | Vibe-IC improves claim discipline by keeping the three axes and hard feasibility separate instead of publishing one collapsed score. ORFS remains ahead in automatic trial execution. |
Eighteen DECLARED_ONLY edges read like eighteen edges nobody got round to wiring. A program now says otherwise. closed_loop_metric_reaches_its_producer asks the question the other two auditors do not: closed_loop_edge_check asks whether a declaration is WELL-FORMED, closed_loop_executable_coverage_check asks whether something ACTUALLY re-enters, and this one asks whether anything COULD.
The predicate is one sentence: a closed-loop edge is a repair, so the step being re-entered has to be able to SEE the quantity its trigger names. If it cannot, re-entering reproduces what it produced before and the loop is inert by construction.
$ python3 programs/closed_loop_metric_reaches_its_producer.py . 21 declared edge(s); REACHABLE=0, UNREACHABLE=2, UNSTATED=19
Both edges that NAME a metric are UNREACHABLE for the same reason: no producer at the fallback step reads it. L19.die_area_budget_um reaches floorplan_contract and a set of checkers and stops there; power__total reaches nothing at step 17 either. The other nineteen name no metric at all, so the question cannot be put to them — and that is reported as a third verdict, UNSTATED, rather than folded into UNREACHABLE. An edge reading "CDC/RDC violation requires RTL change" may well be closeable; merging the two would report 21 unreachable edges and make the flow look worse than it is.
So the page's own note that Vibe-IC has no candidate rewrite executor yet is right — and it blocks more than Phase 4. Closing the area edge 9 → 1 needs three things that do not exist: an L-doc carrying an area constraint the RTL layer reads (today it reaches only the floorplan), an emitter in deterministic_emit_chain that responds to it (none of the six takes an area parameter), and area as a remediable hint kind (the set accepts two, both wiring defects). step_rtl_gen is deterministic, so without all three a re-entry returns byte-identical RTL. The repair loop is a RETRY executor, not a candidate-rewrite executor: it re-runs a deterministic generator and depends on outside state having changed to make the next pass different.
On 2026-08-28 the flow's 68 × 9 gate matrix was probed by injecting, into each of the nine dimensions, the defect its own published question claims to catch. Five reddened; four did not. One finding from that probe changes how the numbers above should be read, and it was re-measured here rather than quoted.
A PPA sign-off gate can keep every structure and simply decide wrong, and nothing notices. The probe switched off a wafer-sort yield verdict with one token — if False and measured + 1e-9 < target: — so a measured 12.5% yield passes a 90% target and the gate exits 0. Every file is still present, every flag still parsed, every dependency edge still declared, every artefact still written. All nine dimension modules were then re-run against that tree: the outcomes are byte-identical to clean, 986 passed / 69 skipped / 8 xfail / 0 failed in both. 0 of the 612 cells change colour. A second injected defect — a sign-off gate that stops reading its signed memo, so the memo says FAIL and the machine record says PASS — is caught by nothing in the repository at all.
It does not weaken what the review above says the framework measures. It narrows what a green count is evidence of: reachable is not bound, and a gate that keeps every structure while deciding wrong is invisible to a structural matrix. The second of those is the one that matters most for PPA, because a PPA verdict IS a decision about a measured number.
Five deterministic gates came out of that probe and are in review, not landed on the pinned source above: every declared step reaches the evaluator; a declared output has a live producer in the source; a disclosure token needs a runtime condition; a red that only means “nothing was there” is counted apart; and a verdict arm decided by a constant. The last of those catches the yield defect at the top of this section — the one shape of the semantic class a source scan can decide. It does not close the class: the memo defect stays open.
ppa_search_run consumes externally produced trial outcomes; without them it emits a plan.ROLLBACK_PROVEN count now reads 2. What remains is EXECUTABLE.Pinned source: vibe-ic v1.12.34 at c51f83082. Direct census, re-derived on that commit: 21 edges over 68 steps; DECLARED_ONLY=18, EXECUTABLE=0, REMEASURED=1, ROLLBACK_PROVEN=2. Dedicated registry: 21 declared edges, 0 BOUND, 21 DECLARED_ONLY, 6 actuators (1 EXECUTABLE), 9 domains (2 EXECUTABLE), 1 controller; registry references verified rc 0, and every EXECUTABLE claim resolves to a file in the tree.
Software-contract scope: 106 selected PPA, closed-loop, wiring, inventory and seven-skill test files — 2,862 passed, 16 skipped and 17 expected failures in 408.15 seconds. One skipped validator-unavailable branch was explicitly reported NOT VERIFIED because the shipped bundled validator made that condition unreachable on this host; it is not counted as proof.
Evidence controls: the generated winner report passed 35 sentences / 139 claims / 9 banned forms; all 14 PPA mutation fixtures discriminated good and bad arms. Cross-layer contracts compared 210 pairs across 21 contracts; end-to-end contracts compared 1,830 pairs across 61 contracts; both corpora reported zero identity conflicts.
These tests establish the software contracts above. They are not a measured three-arm PPA win and do not establish autonomous cross-layer convergence.
This receipt and the metric cards above state the same census, and on 2026-08-28 they did not agree: the cards read REMEASURED 1 / ROLLBACK_PROVEN 2 while this paragraph still read 3 / 0. Nothing read both. A gate now does — page_states_one_figure_twice_check, which compares every figure this page declares on a card against every restatement of it in prose, and which reports this page’s published state as the defect it was written from.
Source pins: OpenROAD 57690eb · LibreLane 33ab648 · OpenROAD-flow-scripts 7386fff. Review scope: the named repositories and revisions, not every EDA flow.