Implementation review · v1.12.34 · 2026-08-28

Closing the PPA loop.

v1.12.34 has real PPA records, parsers, feasibility checks, search contracts and seven PPA skills. Three canonical edges remeasure after an action and two of them are now rollback-proven; no edge is controller-bound yet. A feasibility audit says why: of the 21 declared edges, the two that name a metric cannot reach the producer that would have to change, and 19 name no metric at all. This page separates those facts from the plan.

Historical baseline — v1.11.7

This is the 2026-08-20 snapshot that motivated the five-phase plan, not the current implementation status. At v1.11.7, parsing flow/phase1_phase2_phase3.yaml found 19 closed_loop declarations. The v1.12.34 review below supersedes every present-tense conclusion from this baseline.

4 Performance · loop edges

At v1.11.7: steps 10→7, 20→19, 23→32 and 32→32 were the four performance-related declarations. Their execution-proof levels had not yet been audited.

1 Power · loop edges

At v1.11.7, only IR-drop 24→15 was classified on the power axis; step 33 recorded total power without a declared fallback edge.

0 Area · loop edges

At v1.11.7, the audit found zero area-axis fallback declarations and no declared area metric. Later versions changed this; see the v1.12.34 review.

At that baseline, step 32 declared an aggregator fallback over steps 21–28 after ECO. The declaration was the reason to investigate the lane, not proof that every downstream step was rerun or that rollback worked; those proof levels are measured separately below.

Four programs not referenced by the v1.11.7 flow

Program Flow references Status
ppa_head_to_head_check.py0Gate exists in the hygiene suite, marked run_tolerating_uncheckable, and its fixture pair is recorded as NOT_WRITTEN_YET debt.
ppa_predict_aggregate.py0Pre-synthesis power / area / Fmax estimator. Unreachable from the flow.
ppa_area_threshold_check.py0Synthesises an original / optimised pair with the same recipe and compares cells and wires — exactly the gate the area axis needs.
readme_ppa_extractor.py0Reads the PPA targets the design declares in its own documents.

In that v1.11.7 snapshot, the ppa-predict skill was also referenced zero times. v1.12.34 now gives it a mandatory deterministic aggregator; that is documented in the current review.

What the other two do

Read out of the sources, not from memory: OpenROAD a47e38e, LibreLane bf8cc13, OpenROAD-flow-scripts 3476215. Their PPA is three layers — the actuators inside the tool, the knobs the flow exposes, and a search that drives them.

39Named PPA parameters on LibreLane's four resizer steps (8 + 4 + 13 + 14)
5Search algorithms in ORFS AutoTuner: hyperopt (default), ax, optuna, pbt, random
23Metric checkers LibreLane gates on, including setup, hold, max-slew and max-cap
0Open flows that search across layers — every knob above is place-and-route only

OpenROAD — the actuators

The tool already trades one axis against another, and we use almost none of it.

repair_timing [-setup] [-hold] [-recover_power percent_of_paths_with_slack]
              [-sequence move_list] [-skip_pin_swap] [-skip_gate_cloning]
# -recover_power spends leftover timing slack on power. That is P↔P, in the tool.

set_opt_config [-limit_sizing_area] [-limit_sizing_leakage]
               [-keep_sizing_site] [-keep_sizing_vt]
               [-sizing_area_limit] [-sizing_leakage_limit]
# a ceiling on what fixing timing is allowed to cost in area and in leakage.

replace_arith_modules [-path_count n] [-slack_threshold f] [-target opto_goal]
# swaps arithmetic implementations on critical paths. Micro-architecture, in the tool.

Its metrics are finer-grained than ours too: power splits into internal / switching / leakage / total, and area into stdcell / macros / padcells / cover.

LibreLane — the knobs

Two families matter most, because they are the timing-versus-area trade made adjustable instead of hard-coded.

PL_RESIZER_SETUP_MAX_UTIL_PCT    core area fixing setup is allowed to consume
PL_RESIZER_HOLD_MAX_UTIL_PCT     core area fixing hold is allowed to consume
PL_RESIZER_SETUP_REPAIR_TNS_PCT  what fraction of violating endpoints to even try
# and the same fourteen again, prefixed GRT_, for the post-global-route pass.

It also ships two exploration flows — but read what they actually are:

  • SynthExploration runs nine synthesis strategies in parallel and prints a table for a human to read. It chooses nothing.
  • Optimizing runs three strategies, keeps the smallest design__instance__area, then tries FP_CORE_UTIL=99 and falls back to 40 on failure. Its own header calls it “a custom demo flow”, and it optimises a single objective.

ORFS AutoTuner — the search

This is the part worth stealing. Ray Tune drives it, and the objective is explicitly a head-to-head against a reference run.

coeff_perform, coeff_power, coeff_area = 10000, 100, 100

eff_clk_period = clk_period - worst_slack        # only when slack is negative

performance = percent(eff_clk_period_ref, eff_clk_period)
power       = percent(reference["total_power"], metrics["total_power"])
area        = percent(100 - reference["final_util"], 100 - metrics["final_util"])

ppa   = performance*10000 + power*100 + area*100
score = (upper_bound - ppa) * (step/100)**-1 + (ppa/10) * num_drc

Four decisions in there are worth naming:

  1. Everything is a percentage improvement over a baseline, never an absolute number. The tuner is a head-to-head machine.
  2. Performance is weighted 100× power and area. That is a value judgement about a market, not a fact — and it is hard-coded.
  3. DRC count is a penalty and step is a divisor, so a run that dies early or leaves violations cannot buy a good score.
  4. The clock target itself is inside the search space (_SDC_CLK_PERIOD).

The space it searches, per design:

_SDC_CLK_PERIOD                     float [3.5, 7.0]
CORE_UTILIZATION                    int   [20, 50]
CORE_ASPECT_RATIO                   float [0.5, 2.0]
CELL_PAD_IN_SITES_GLOBAL_PLACEMENT  int   [0, 3]
CELL_PAD_IN_SITES_DETAIL_PLACEMENT  int   [0, 3]
PLACE_DENSITY_LB_ADDON              float [0.0, 0.2]
CTS_CLUSTER_SIZE                    int   [10, 200]
CTS_CLUSTER_DIAMETER                int   [20, 400]
_FR_LAYER_ADJUST                    float [0.1, 0.3]
The gap found in the pinned survey

In the pinned OpenROAD, LibreLane and ORFS source revisions listed below, the exposed tuner knobs are place-and-route knobs. AutoTuner does not touch synthesis strategy, RTL or micro-architecture; LibreLane's synthesis exploration uses nine discrete presets and does not feed placement. None of those three pinned flows performs a unified cross-layer search.

The plan

Five phases. Each one is useful on its own and each one is a precondition for the next, which is why they are numbered rather than listed. Every phase closes on a command that returns zero only when the work is really on main.

PHASE 0 · 1 worker

Measure the two axes we are blind on

PowerArea

Nothing downstream can converge on a metric the flow does not record. Add the metrics as declared required_outputs on the steps that already produce them, using OpenROAD's own names so a reader can cross-check us against ORFS without translation.

step  9  design__instance__area, __count
step 17  design__instance__utilization
         design__core__area
step 21  route__wirelength
step 33  power__internal|switching
              |leakage|total  (split)
step 37  design__die__area

Every unmeasured field is the literal NOT_MEASURED. A plausible default here would hide precisely the hole this phase exists to expose.

CLOSES · flow_metric_coverage_check --axis area --axis power → exit 0
PHASE 1 · 1 worker

Give power and area a fallback edge

PowerArea

This is the phase the whole plan is named after. Two new closed_loop declarations, in the same shape the timing edges already use — plus a ceiling on what fixing timing may cost, which is the trade LibreLane exposes and we do not.

step 33  closed_loop: fallback_to: 17
  trigger: power__total above the
           design's declared budget

step  9  gate: ppa_area_threshold_check
  trigger: area above the ceiling

step 32  + RESIZER_SETUP_MAX_UTIL_PCT
         + RESIZER_HOLD_MAX_UTIL_PCT
         + repair_timing -recover_power

ppa_area_threshold_check.py already synthesises an original / optimised pair and compares cells and wires. It needs wiring, not writing.

CLOSES · matrix_mutation_ledger --replay D5-PHANTOM-EDGE --step 33 → REDDENED
PHASE 2 · 1 worker

Wire the head-to-head, and pay its fixture debt

PerformancePowerArea

ppa_head_to_head_check.py opens by stating the problem better than we can: our published numbers prove an AI can produce relatively correct RTL; they do not prove it can produce better silicon, and a reviewer is entitled to answer them with “so what?”

arm A  vibe-ic     our flow, our defaults
arm B  LibreLane   Classic, its defaults
arm C  ORFS        + AutoTuner, tuned best
# arm C is the one that makes the
# comparison worth reading.

Beating an untuned baseline proves nothing. The claim only means something against a flow that was allowed to search. What is missing is the fixture pair, a flow step that calls the program, and the third arm.

CLOSES · gate_fixtures/ppa_head_to_head_records.py on main AND named in the flow
PHASE 3 · 6 workers

The search layer — and the fleet is the advantage

PerformancePowerArea

AutoTuner's bottleneck is that a good search wants hundreds of full flow runs; they solve it with a Ray cluster. We already have a fleet and a dispatch discipline. That is a structural fit, not a coincidence.

Port the objective, with one deliberate change: the weights become a declared input, not a constant.

# ORFS hard-codes 10000/100/100. A 100:1 preference for speed over power is a
# value judgement about a market, and it belongs to the design, not to us.
ppa_weights:
  performance: <from L19, no default>
  power:       <from L19, no default>
  area:        <from L19, no default>
# A design that declares no preference inherits ORFS's ratio EXPLICITLY, recorded
# in the report as "inherited, not chosen", so a reader can see whose call it was.

Keep the two anti-cheating terms verbatim, because they are the reason the score cannot be gamed: num_drc as a penalty, and (step/100)**-1 so a run that stops early is scored on how far it got.

CLOSES · a ≥50-configuration search beats our own default run, all 50 records published
PHASE 4 · 6 workers

Search across layers — beyond the pinned tuners

PerformancePowerArea

Everything in Phase 3 is a place-and-route knob, because that is all the open flows tune. The search space we can offer that they cannot:

LayerLeverWho searches it today
Place & routeutilisation, density, padding, CTS clusteringORFS AutoTuner
SynthesisSYNTH_STRATEGYLibreLane — 9 fixed presets, a human picks
Arithmeticadder / multiplier architectureOpenROAD, narrowly — replace_arith_modules
RTLpipelining, resource sharing, encodingnone in the pinned survey
Micro-architecturethe structure the spec never pinnednone in the pinned survey

The bottom two rows motivate the plan. An agent can propose RTL and micro-architecture changes, but an accepted rewrite must be required to pass equivalence, lint and STA and to preserve their verdicts. v1.12.34 does not yet prove that this gate set executes for every cross-layer candidate.

CLOSES · cross-layer search beats the best PnR-only search, winning RTL passes step 13

Discipline

A PPA number is a claim about silicon. These are the ways such a claim goes wrong, and every one of them has precedent in this repository.

  • No relaxing to win. Never hand-edit a GDS, delete violating geometry, move a pin, or loosen a rule deck. A score obtained that way is worth less than the failure it replaces.
  • One run tree per report. Every figure in a head-to-head resolves from one run, and finding two is a refusal, not a choice of the newer. Carry a sha256 or a die bbox beside any number quoted across a boundary — on 2026-08-20 one design read as both passing and failing because two different artefacts wore the same name.
  • Unmeasured is not clean. A metric the run never produced is NOT_MEASURED and blocks. “We did not look” and “we looked and it was fine” must not produce the same artefact.
  • A tuned arm must be allowed to tune. If arm C is ORFS, it gets its AutoTuner budget. Publishing a win over a deliberately weakened opponent is the same defect as a gate that cannot fail.
  • Every new gate ships mutation-proved. A green cell with no reachable red is a certificate, not a measurement — which is this repository's own rule, and it applies to these gates too.

What this plan does not do

  • It does not replace step 32. The timing loop works; it gains an area ceiling and nothing else.
  • It does not fork a tuner. Phases 3 and 4 call ORFS and LibreLane as arms, and our own search reuses the fleet dispatch that already exists.
  • It does not touch 0.5ic/d3, which is red for an unrelated reason: step 0.5ic is newer than every published cell, so no published run has produced its declared outputs.
  • It sets no numeric target. A goal like “10% better area” before the first honest three-way measurement would be a threshold the design specification never declared.

Sources read for this plan: flow/phase1_phase2_phase3.yaml @ v1.11.7 · OpenROAD a47e38e · librelane bf8cc13 · OpenROAD-flow-scripts 3476215. Every count in the first section was produced by parsing the flow, not recalled.

IMPLEMENTATION REVIEW · 2026-08-28

What is real now — v1.12.34

The page above is the v1.11.7 plan written on 2026-08-20. This appended review supersedes its present-tense status claims; it does not rewrite that historical plan.

21declared closed-loop edges
68canonical flow steps
18DECLARED_ONLY
0EXECUTABLE
1REMEASURED
2ROLLBACK_PROVEN
1301of 1301 top-level programs reachable — but only 93 sit on a clause that can fail a step
10of 10 PPA checkers have machine runners
REVIEW VERDICT · NEEDS ATTENTION

Stronger wiring and claim control; still not an autonomous optimizer

v1.12.34 keeps the canonical PPA records, scoped identities, provenance, hard feasibility, Pareto comparison, search budgets, closure state machines, five report parsers and seven PPA skills. It adds a more important property around them: every top-level program is now reachable, all ten checker-shaped PPA programs have machine runners, and the generated PPA report has a claims file plus a gate that refuses prose stronger than its evidence.

Reachable is not the same as bound. The dedicated registry still reports 21 declared edges, zero BOUND controllers and 21 DECLARED_ONLY edges; its one executable controller is bound to no canonical edge. The generic search still records externally produced trials instead of launching a complete cross-layer campaign itself.

REMEASURED means that an action ran and the relevant condition was measured again. It does not by itself mean converged, accepted, shipped or rollback-proven. The tiers are nested, so an edge counted at a higher tier is no longer counted at a lower one: 4→1 (functional simulation) is the one REMEASURED edge, and the two PPA-related ones, 23→32 and 32→32, have since been promoted to ROLLBACK_PROVEN.

What v1.12.34 actually changed

The release closes reproducibility and wiring gaps around PPA. It does not promote declaration into convergence.

Areav1.12.34 contractHonest boundary
Program reachabilityThe wiring audit covers 1,301 top-level programs and 657 checker-shaped ones; all ten ppa_* checkers resolve to a machine runner — eight from the hygiene lane, one (head-to-head) also from the flow yaml, and two from another program.A separate automatic-verdict audit still lists 27 non-PPA gates outside an automatic verdict path. Reachability is not proof that every program is a canonical flow gate.
ppa-diagnoseThe skill now runs ppa_diagnostic_router.py first, continues only on an explicit HANDOFF, then builds a path-and-hash evidence context with ppa_agent_context_build.py.PROGRAM_DECIDED, REFUSED and UNDETERMINED stop the AI path. Missing evidence may not be replaced by a fluent hypothesis.
ppa-measureThe skill now requires ppa_signoff_records and ppa_eco_spare_records to produce canonical records; ppa_report_gen then produces report.md plus claims.json, and ppa_page_claim_check checks citations, numbers and evidence strength.The two record producers are operator-run today, not canonical flow clauses: spare_preservation.json is conditional and undeclared, so making it an unconditional dependency would falsely red every design without spares.
ppa-predictThe skill now makes ppa_predict_aggregate.py own the cell-count-to-area/power/Fmax arithmetic, PDK constants, ranges, status labels and assumptions.A missing or zero cell-count anchor is rc 2 CANNOT CHECK. The result stays ESTIMATED / RTL_PROXY and is explicitly ineligible for physical-PPA claims.
Published-report claimsThe repository's winner report passes over 35 sentences, 139 claims and 9 banned unqualified forms. PPA mutation fixtures prove the good arm passes and a promoted or broken claim fails.That gate protects the generated repository report/claims pair, not this static marketing page. This page is instead pinned to the source commit and verification receipt below.

Layers and flow edges

The vocabulary reaches beyond place and route, but declaration, execution and proof are separate levels.

LayerCurrent lever surfaceRelevant flow edgesMeasured status
RTLpipelining and state encoding; specification permission and equivalence mode are recorded2→1 · 3→1 · 4→1 · 5→14→1 REMEASURED; the others DECLARED_ONLY
Synthesis and constraintsAREA 0..3 and DELAY 0..4 strategy vocabulary8→7 · 9→1 · 10→7 · 13→9 · 14→9DECLARED_ONLY
Arithmeticadder and multiplier architecture choices are declared; Vibe-IC has no candidate rewrite executor yetcross-layer search space, not a bound canonical edgeDECLARED_ONLY
Microarchitecturemodule hierarchy only; not general microarchitecture explorationcross-layer search space, not a bound canonical edgeDECLARED_ONLY
Analogpost-extraction and hardware-correlation fallback declarationsA7→A3 · A9→A3DECLARED_ONLY
Physical design and sign-offfloorplan, placement, ECO spares, CTS, routing and constraint vocabulary20→19 · 23→32 · 24→15 · 25/26/27→21 · 28→15 · 31→32 · 32→32 · 33→1723→32 and 32→32 ROLLBACK_PROVEN — the undo has a named test

What was learned from the open flows

SystemWhat it already doesVibe-IC's position
OpenROADThe real implementation actuator: timing repair, buffering, sizing, VT swaps and narrowly scoped arithmetic replacement. Arithmetic replacement currently implements setup optimization; hold, power and area targets remain unavailable.Vibe-IC does not replace OpenROAD's optimizer; it adds evidence identity, provenance, feasibility and controlled adoption around it.
LibreLaneSynthesisExploration runs nine synthesis presets and reports them for a human choice. Its separate Optimizing demo selects the smallest-area result from three presets, then tries high utilisation with a lower-utilisation fallback.Vibe-IC has stricter scope and evidence contracts, but has not yet connected them to an equally complete trial executor.
ORFS AutoTunerActually launches and manages EDA trials through Ray Tune. The current CLI exposes hyperopt, ax, optuna, pbt and random, with reference-relative PPA terms plus early-stop and DRC penalties.Vibe-IC improves claim discipline by keeping the three axes and hard feasibility separate instead of publishing one collapsed score. ORFS remains ahead in automatic trial execution.

EXECUTABLE = 0 is not a wiring backlog

Eighteen DECLARED_ONLY edges read like eighteen edges nobody got round to wiring. A program now says otherwise. closed_loop_metric_reaches_its_producer asks the question the other two auditors do not: closed_loop_edge_check asks whether a declaration is WELL-FORMED, closed_loop_executable_coverage_check asks whether something ACTUALLY re-enters, and this one asks whether anything COULD.

The predicate is one sentence: a closed-loop edge is a repair, so the step being re-entered has to be able to SEE the quantity its trigger names. If it cannot, re-entering reproduces what it produced before and the loop is inert by construction.

$ python3 programs/closed_loop_metric_reaches_its_producer.py .
21 declared edge(s); REACHABLE=0, UNREACHABLE=2, UNSTATED=19

Both edges that NAME a metric are UNREACHABLE for the same reason: no producer at the fallback step reads it. L19.die_area_budget_um reaches floorplan_contract and a set of checkers and stops there; power__total reaches nothing at step 17 either. The other nineteen name no metric at all, so the question cannot be put to them — and that is reported as a third verdict, UNSTATED, rather than folded into UNREACHABLE. An edge reading "CDC/RDC violation requires RTL change" may well be closeable; merging the two would report 21 unreachable edges and make the flow look worse than it is.

So the page's own note that Vibe-IC has no candidate rewrite executor yet is right — and it blocks more than Phase 4. Closing the area edge 9 → 1 needs three things that do not exist: an L-doc carrying an area constraint the RTL layer reads (today it reaches only the floorplan), an emitter in deterministic_emit_chain that responds to it (none of the six takes an area parameter), and area as a remediable hint kind (the set accepts two, both wiring defects). step_rtl_gen is deterministic, so without all three a re-entry returns byte-identical RTL. The repair loop is a RETRY executor, not a candidate-rewrite executor: it re-runs a deterministic generator and depends on outside state having changed to make the next pass different.

What a probe found that a green ledger could not

On 2026-08-28 the flow's 68 × 9 gate matrix was probed by injecting, into each of the nine dimensions, the defect its own published question claims to catch. Five reddened; four did not. One finding from that probe changes how the numbers above should be read, and it was re-measured here rather than quoted.

A PPA sign-off gate can keep every structure and simply decide wrong, and nothing notices. The probe switched off a wafer-sort yield verdict with one token — if False and measured + 1e-9 < target: — so a measured 12.5% yield passes a 90% target and the gate exits 0. Every file is still present, every flag still parsed, every dependency edge still declared, every artefact still written. All nine dimension modules were then re-run against that tree: the outcomes are byte-identical to clean, 986 passed / 69 skipped / 8 xfail / 0 failed in both. 0 of the 612 cells change colour. A second injected defect — a sign-off gate that stops reading its signed memo, so the memo says FAIL and the machine record says PASS — is caught by nothing in the repository at all.

It does not weaken what the review above says the framework measures. It narrows what a green count is evidence of: reachable is not bound, and a gate that keeps every structure while deciding wrong is invisible to a structural matrix. The second of those is the one that matters most for PPA, because a PPA verdict IS a decision about a measured number.

Five deterministic gates came out of that probe and are in review, not landed on the pinned source above: every declared step reaches the evaluator; a declared output has a live producer in the source; a disclosure token needs a runtime condition; a red that only means “nothing was there” is counted apart; and a verdict arm decided by a constant. The last of those catches the yield defect at the top of this section — the one shape of the semantic class a source scan can decide. It does not close the class: the memo defect stays open.

What v1.12.34 still does not close

  1. Bind a real edge. The current registry reports 21 declared edges, zero bound.
  2. Run candidates, not just manifests. ppa_search_run consumes externally produced trial outcomes; without them it emits a plan.
  3. Ship the repaired implementation. Step 32 currently runs after GDS, DRC and LVS; it lists the invalidated downstream steps but does not rerun them, and its ECO netlist is not the shipped implementation.
  4. Rollback is proven. A named test exercises the revert decision, and the measured ROLLBACK_PROVEN count now reads 2. What remains is EXECUTABLE.
VERIFICATION RECEIPT

What this review actually ran

Pinned source: vibe-ic v1.12.34 at c51f83082. Direct census, re-derived on that commit: 21 edges over 68 steps; DECLARED_ONLY=18, EXECUTABLE=0, REMEASURED=1, ROLLBACK_PROVEN=2. Dedicated registry: 21 declared edges, 0 BOUND, 21 DECLARED_ONLY, 6 actuators (1 EXECUTABLE), 9 domains (2 EXECUTABLE), 1 controller; registry references verified rc 0, and every EXECUTABLE claim resolves to a file in the tree.

Software-contract scope: 106 selected PPA, closed-loop, wiring, inventory and seven-skill test files — 2,862 passed, 16 skipped and 17 expected failures in 408.15 seconds. One skipped validator-unavailable branch was explicitly reported NOT VERIFIED because the shipped bundled validator made that condition unreachable on this host; it is not counted as proof.

Evidence controls: the generated winner report passed 35 sentences / 139 claims / 9 banned forms; all 14 PPA mutation fixtures discriminated good and bad arms. Cross-layer contracts compared 210 pairs across 21 contracts; end-to-end contracts compared 1,830 pairs across 61 contracts; both corpora reported zero identity conflicts.

These tests establish the software contracts above. They are not a measured three-arm PPA win and do not establish autonomous cross-layer convergence.

This receipt and the metric cards above state the same census, and on 2026-08-28 they did not agree: the cards read REMEASURED 1 / ROLLBACK_PROVEN 2 while this paragraph still read 3 / 0. Nothing read both. A gate now does — page_states_one_figure_twice_check, which compares every figure this page declares on a card against every restatement of it in prose, and which reports this page’s published state as the defect it was written from.

Source pins: OpenROAD 57690eb · LibreLane 33ab648 · OpenROAD-flow-scripts 7386fff. Review scope: the named repositories and revisions, not every EDA flow.