# Fifteen prompts, five criteria: what a typed registry actually changed

> I couldn't run this on the registry I built at work. So I rebuilt the pattern on a component set I still have, ran 90 generations, and scored them.

- Author: Alejandro Haydar
- Published: 2026-08-19
- Taxonomy: Design Systems, Design Engineering, AI, Workflows, Showcase
- Canonical: https://alejandroastroport.netlify.app/showcases/registry-eval/

---
import CodeBlock from "../../components/mdx/CodeBlock";
import resultsSvg from "../../assets/images/registry-eval/results.svg?url";

I have a showcase on this site about a typed component registry I built at Thought Industries — a file that hands coding agents a closed world so they stop inventing components that don't exist. At the bottom of it, in the results section, sat this sentence:

> Agent behaviour improved in tested Cursor sessions. This is an observed pattern, not a controlled metric yet.

That's an honest sentence and a weak one. It says *I watched some things happen.* Anyone can write it. So I ran the measurement.

**One line of housekeeping, then I'll move on:** 90 generations, one model, 15 prompts, and — importantly — not the Thought Industries registry, because I no longer have access to that repo. What I tested is the pattern, rebuilt over a component library I do have.

---

## What I could test, and what I couldn't

I left Thought Industries. The registry went with the codebase.

The temptation was to write "here's how I *would* evaluate this" — a methodology page, no numbers. That's the artifact of someone who has never run an eval, and it's indistinguishable from one.

The claim under test was never *tiOS is good.* It was **a typed closed world changes what an agent generates.** That's testable against any bounded component set. This site runs on one: 46 Subframe-synced components with real variant enums, compound children, and a semantic token file.

So the registry got **derived from source**, not written by me:

<CodeBlock
  client:visible
  filename="extract-registry.mjs — output"
  hideLineNumbers
  code={`atoms: 46   with variant enums: 13   compound: 23   tokens: 68`}
/>

This turns out to be a stronger position than the hand-authored version I'd have tested at work. Nobody can argue I tuned the registry to flatter the eval, because I didn't write it. A script read the components that exist and enumerated them. If the extractor is wrong, the registry is wrong in the same direction for every condition.

---

## The design

Three conditions, identical task framing, one variable — the context block:

| Arm | What the model got |
|---|---|
| **none** | "Generate React UI code." Nothing else. |
| **prose** | A paragraph: ShadCN-style components at `@/subframe/components`, semantic tokens, don't invent primitives. |
| **registry** | The full enumerated inventory — every component, every variant enum, every legal compound child, every token. |

The middle arm is the one that matters, and it wasn't in my first sketch of this. Without it, a score that goes up tells you *context helps* — not *the enumeration helps*. Those are different claims and only one of them is about the work I did.

Five criteria. Four scored by script against the derived registry, so there's no LLM judge and no rubric drift: imports resolve to real components, styling uses semantic tokens, composition is legal, variant values exist in the enum. The fifth — did it actually do the job — needs a human.

Fifteen prompts, stratified rather than freehanded, because prompts written by the person who built the thing will quietly target its strengths: five single-component, five composition-requiring, five **traps** that ask for something the library doesn't have (a card, a popover, a form, pagination, a hover card — all fluent ShadCN, none of them present).

Generation ran through the CLI with every tool disabled, in an empty directory. That part is not fussiness: if the model can read files, the no-context arm quietly stops being a no-context arm.

---

## What came back

<figure class="full-width my-8">
  <img src={resultsSvg} alt="Results across five criteria and three conditions: the registry arm scores 100% on all four mechanical criteria; the prose arm scores 43% on imports resolving, below the no-context baseline; all three arms complete the task at near-ceiling rates." class="w-full h-auto" />
</figure>

<div class="breakout my-8">

|  | none | prose | registry |
|---|---|---|---|
| C0 used the design system | 0% | 100% | 100% |
| C1 imports resolve | 97%※ | **43%** | 100% |
| C2 semantic tokens only | 10% | 100% | 100% |
| C3 legal composition | 100%※ | 97% | 100% |
| C4 variants in enum | 97%※ | 93% | 100% |
| **on-system and clean** | **0%** | **40%** | **100%** |

</div>

※ Vacuous. The no-context arm hand-rolls raw HTML with inline styles — it imported nothing, so it had nothing to get wrong.

The prose-versus-registry gap on imports resolving (13/30 against 30/30) gives a Fisher's exact **p ≈ 6×10⁻⁷**. On the combined on-system-and-clean measure, **p ≈ 2×10⁻⁷**. Small eval, large effect — the separation between a paragraph and an enumeration is not a sampling artefact.

That asterisk was the first real thing the harness taught me. Scored naively, the arm with no design-system knowledge at all "wins" three of four criteria, because every criterion is phrased as an absence of violations and you cannot violate a system you never touched. The fix was a separate adoption check and a headline metric of *on-system **and** clean*. I'd have shipped the flattering-to-nobody version of this table if I hadn't looked at why the numbers were strange.

---

## The result I didn't expect

The prose arm is **worse than no context at all** on hallucinated imports.

Not worse than the registry — worse than the baseline. Same prompt, both arms told to use the design system, one told what's in it:

<div class="breakout my-8">

<CodeBlock
  client:visible
  language="tsx"
  filename="prose arm — B3, confirmation dialog"
  highlightedLines={[3,4,5,6,7]}
  code={`import {
  Dialog,
  DialogContent,      // none of these
  DialogHeader,       // exist in
  DialogTitle,        // this
  DialogDescription,  // library
  DialogFooter,
  Button,
} from "@/subframe/components";`}
/>

<CodeBlock
  client:visible
  language="tsx"
  filename="registry arm — B3, same prompt"
  highlightedLines={[2,5]}
  code={`import React from "react";
import { Dialog } from "@/subframe/components/Dialog";
import { Button } from "@/subframe/components/Button";

<Dialog open={open} onOpenChange={onOpenChange}>
  <Dialog.Content>`}
/>

</div>

This library uses compound children: `Dialog.Content`, `Table.Row`, `Select.Item`. The model, told a real import path but not the inventory, wrote the flat ShadCN names it has seen ten thousand times — `DialogContent`, `TabsList`, `SelectTrigger`, `Card`, `Text`. Confident, fluent, and not there.

The paragraph didn't reduce guessing. It gave the guesses a real address to arrive at. The no-context arm never made this mistake because it never tried to import anything — it hand-rolled a div.

So the control arm earned its cost. Prose closed the token gap completely (10% → 100%): a sentence is enough to stop `bg-blue-500`. Prose could not close the inventory gap at all. **Only the enumeration did that**, and that's the specific claim the tiOS page couldn't previously make.

---

## The criterion that could have gone against me

Conformance criteria can only flatter a registry. A generation that emits one `<Button>` for every prompt scores four out of four — real component, real token, legal nesting, real variant — while doing nothing. So the fifth criterion asks the only question the script can't: **did it actually do the job?**

That one needs a human, and it needs the human not to know which arm they're grading. I scored 27 samples blind — the trap and composition prompts, nine per arm, shuffled and tagged with hashes, with the key closed until the verdicts were written down.

| | did the job |
|---|---|
| none | 9/9 |
| prose | 9/9 |
| registry | **8/9** |

Fidelity sat at the ceiling everywhere, and the single incomplete result is in the registry arm — a pagination control that rendered previous and next as bare chevrons instead of labelled actions. I could construct an argument that it deserved a pass. I'm not going to, because I only noticed the argument after unblinding told me which arm it belonged to, and revising a blind judgement at that point is just bias with extra steps.

Two things follow. The registry did **not** buy its conformance score by refusing to build things or by quietly under-delivering — the failure mode I built this criterion to catch didn't happen. And with nine per arm, this measure could only ever have caught a catastrophe: it takes a 56-point gap to reach significance at that size, so 9/9 against 8/9 means *nothing broke*, not *the registry wins here*.

The more interesting reading is what fidelity being flat across all three arms implies. **Every condition solved the user's problem.** The no-context arm built a working validating sign-up form, a working hover card, a working team table — all of it hand-rolled, off-system, styled with `bg-gray-50`, and all of it functional. What the registry changed was never whether the model could do the task. It was whether the result belonged in the codebase.

The gap isn't capability. It's compliance.

---

## What this doesn't prove

The registry arm scored 100% on all four mechanical criteria. I don't trust a perfect score, and neither should you:

- **The criteria saturated.** They separate registry from prose cleanly. They cannot discriminate *within* the registry arm at all — every generation passed. A second version needs harder criteria: accessibility, state handling, whether the composition is idiomatic rather than merely legal.
- **I fixed two scorer bugs after seeing results, and the registry score went up.** That is the exact shape of motivated reasoning, so: arbitrary *layout* values like `max-w-[480px]` were being counted as token violations against a rule I'd written saying layout utilities are fine, and `React.Fragment` was counted as illegal composition. Both were false failures against the registry arm. C2 moved from 80% to 100%. The pre-fix numbers are in git history.
- **Fidelity is a guardrail, not a finding.** Nine samples per arm, trap and composition strata only. It can detect the registry arm collapsing into refusals. It cannot resolve anything subtler, and I make no claim from 9/9 versus 8/9.
- **One model, one library, 15 prompts.** Everything here is a claim about `claude-sonnet-5` generating against one Subframe component set.
- **This is not the Thought Industries registry.** Same pattern, different closed world.

---

## What I'd keep

The harness is four files and no framework: a script that derives the registry from source, a prompts file, a runner, a scorer. It's boring on purpose — it's evidence, not a product. Anyone can clone the repo and re-run it, which is the only property that makes a number worth publishing.

The finding I'd actually carry into the next team is the middle arm. "Tell the agent about your design system" is the advice everyone gives, and in this sample it made one failure mode worse, not better — because natural language invites the model to fill gaps and an enumeration doesn't have gaps to fill. The registry works for a duller reason than it seems: not because it explains the system, but because it closes it.

That's a claim I can now point at a table for.