TL;DR: GUML is a tiny, indentation-based UI language designed to be written by a language model, and a Rust compiler that turns it into React, Svelte 5, static HTML, Web Components, a JSON UI tree, A2UI or MCP-UI. The model writes 24 lines. The compiler writes the fetch, the abort, the retry with backoff, the stale-while-revalidate cache, the optimistic update, the rollback and the accessible names: 253 lines of React for the task app below. Every error comes back in one pass, with a code and a suggestion, and the obvious ones are fixed without asking the model again.
cargo install guml-cli # the `guml` command pnpm add @guml/core # the same compiler as WebAssembly pip install guml # and from Python🔗 GitHub · 📦 @guml/core on npm · 📄 Research report
Table of Contents
- Why I built this
- The 24-line argument
- A tour of the language
- Inside the compiler
- A lexer that cannot decide alone
- Every error in one pass
- Did you mean card?
- Repair without a model
- What 24 lines become
- Seven backends, one vocabulary
- Themes and capabilities
- What is measured, and what is not
- A real model tried it
- The playground and the chat
- Proof: what I ran for this post
- What I found while writing this
- The question still open
- What I learned
- What's next
- Try it
Why I built this
Ask any AI app builder for a task list and watch what comes back: two hundred lines of React. useState, useEffect, an AbortController, a loading flag, an error flag, an optimistic update, a rollback, a className string per element, and aria-labels if you're lucky.
Almost none of that is a decision. It's convention: the same fetch-with-cancellation shape, the same rollback, the same focus ring. The model is being paid, in tokens, in latency and in error probability, to re-derive a compiler's output by hand, every single time.
Tailwind makes the point sharply. className="mt-6 flex items-center justify-center gap-3" is about fourteen tokens that express three tokens of intent: row, centered, normal gap.
So the question behind GUML is simple: what if the model only wrote the decisions, and a compiler wrote the conventions? There's no widely used representation that sits between "natural-language prompt" and "framework source code". Every AI builder jumps the whole distance in one hop. GUML is the stop in the middle.
flowchart LR P["prompt"] --> M["LLM"] M -->|"24 lines, ~185 tokens"| G["GUML source"] G --> C["guml compiler (Rust)"] C -->|"253 lines"| R["React + TS"] C --> S["Svelte 5"] C --> H["static HTML"] C --> W["Web Components"] C --> J["JSON UI tree / A2UI / MCP-UI"] C -.-> X["fetch, abort, retry, cache,<br/>optimistic apply, rollback,<br/>loading and empty states, ARIA"]
The dashed box is the whole idea. Nothing in it appears in the source the model writes, so the model can't spend tokens on it, and it can't get it wrong.
The 24-line argument
This is the task-list fixture, fixtures/b.guml, complete:
page Tasks
type Task {id, title, done:bool, createdAt:date}
data tasks:Task[] GET /api/tasks
add POST /api/tasks {title} optimistic:prepend
save PATCH /api/tasks/{id} {done} optimistic
drop DELETE /api/tasks/{id} optimistic
state draft=""
state filter=all|open|done
head Tasks — {tasks.open.count} open
form >tasks.add{title:draft}; draft=""
input draft aria="New task" placeholder="Add a task…"
btn Add primary disabled={!draft.trim()} busy="Adding…"
tabs filter
list tasks where={filter}
check {done} >tasks.save
text {title} strike={done}
btn Delete quiet aria="Delete {title}" >tasks.drop
empty Nothing here yet.
Read it once and you know the whole app: a typed resource with three mutations (two of them optimistic), a text draft, a three-way filter, a live count in the heading, a form, tabs, and a list with an empty state.
Here it is compiled in the browser. This is the docs playground running the Rust compiler as WebAssembly:

Notice the small line under the input: tabs — not lowered by this compiler version. The in-browser preview renderer doesn't draw tabs yet, and it says so instead of drawing something approximate. That's one of the project's four invariants, and it'll come up again.
A tour of the language
The whole syntax fits in a paragraph, which is the point: the full spec plus the vocabulary costs a model about 2,935 estimated tokens of prompt, and the spec file states a budget of 3,000.
Lines are structure. Indentation is nesting, 2 spaces per level. No closing tags, no braces for blocks, no imports. A tab is an error (GUML0001).
Directives declare the data model:
page Counter title="Clicks" // name -> component
type Task {id, title, done:bool} // fields default to string
state count=0 // type inferred from the initial value
state filter=all|open|done // an enumerated domain; first value is initial
data tasks:Task[] GET /api/tasks // a resource
add POST /api/tasks {title} optimistic:prepend
on {filter} >tasks.list // an effect: re-runs when filter changes
Elements are tag [positionals] [name=value] [>action]. Positionals can be a label, a modifier, a {binding}, a /route or an #anchor. The action comes last because > swallows the rest of the line.
Bindings are read-only expressions: paths, comparison, boolean logic, arithmetic, and a fixed set of aggregates (.count .sum .open .done .trim .lower .upper). {tasks.open.count} is "how many rows have the bool field unset". It works on invoices too, where the bool is paid, because the compiler reads the flag from the row type rather than from a name it hoped for.
Actions are statements separated by ;: >count++, >count=0, >tasks.add{title:draft}; draft="".
The vocabulary is closed. There are 49 primitives in a registry (card row col section nav hero form tabs list table btn input select metric modal toast…), 24 modifiers (primary ghost quiet danger sm lg center between…), and 12 global attributes. There is no class attribute: presentation belongs to the theme.
def declares a user component. It's a compile-time macro with parameters and one slot, and by the time codegen runs there's no trace of it.
Two conformance levels. core is markup only: no I/O, no state, meant to be safe to render from an untrusted agent. app adds state, data, actions and the repeaters. An app construct at the core level is GUML0091.
Escape hatches. js and raw blocks pass through verbatim and unchecked. Each one emits a GUML0090 note, so the escape-hatch rate stays countable. That rate is one of the things the research measures.
Inside the compiler
The compiler is a Rust workspace of 14 crates, about 33,000 lines. The heaviest are guml-compiler (about 10.7k lines, the analysis passes) and guml-codegen (about 9.1k lines, the backends).

The driver is short enough to quote. From crates/guml-compiler/src/lib.rs:
let parsed = guml_parser::parse(src, reg);
// ...
expand::expand(&mut program, &mut diagnostics); // def macros
sema::analyse(&program, reg, &mut diagnostics); // scope + accessibility
validate::validate(&program, reg, &mut diagnostics); // structure, methods, unused
types::check(&program, &mut diagnostics); // expression types, aggregates
flowchart TD
SRC["GUML source"] --> LEX["lex: line-oriented, never fails"]
LEX --> PARSE["parse: registry-aware, recovers"]
PARSE --> EXP["expand: def macros, slots"]
EXP --> SEMA["sema: bindings, scope, accessible names"]
SEMA --> VAL["validate: mutations, methods, unused state"]
VAL --> TY["types: Num / Str / Bool / List / Unknown"]
TY --> GATE{"any errors?"}
GATE -- "yes" --> DIAG["every diagnostic, one pass"]
GATE -- "no" --> CG["codegen"]
CG --> R["react"]
CG --> SV["svelte"]
CG --> HT["html"]
CG --> WC["wc"]
CG --> JS["json"]
CG --> A2["a2ui"]
CG --> MC["mcp-ui"]All the analysis passes always run, even after an error, so diagnostics accumulate. Codegen runs only on an error-free program. There's no separate IR: the AST, guml_ast::Program, is what every backend consumes, and each backend lowers expressions through one shared function.
The type checker is the newest stage and, according to the docs, the one that paid for itself most directly. It found a live bug in a published example where {invoices.open.count} had been compiling to a filter on a field the type didn't declare: always zero, and no test noticed.
A lexer that cannot decide alone
GUML has an ambiguity no lexer can resolve on its own. The comment at the top of crates/guml-syntax/src/lib.rs puts it in two lines:
btn Decrement ghost >count-- // Decrement is a label, ghost is a modifier
p Press the buttons to change. // the whole remainder is prose content
Same shape, opposite meanings. Whether the rest of a line is structure or prose depends on the tag, and the tag's kind lives in the registry. So the lexer produces both readings of every line: a token list and the raw text. Then the parser, which knows the registry, picks one.
The rule is subtle: a text tag stays prose unless the line contains an = whose name the registry accepts for that tag. So p Set x=1 to enable is prose, and text {title} strike={done} is structure. Step through it:
const C = {
tag: [247, 59, 32], label: [235, 235, 235], mod: [100, 181, 246], bind: [181, 131, 224],
attr: [42, 157, 143], action: [255, 209, 102], prose: [150, 150, 150]
};
const NAMES = { tag: 'tag', label: 'label', mod: 'modifier', bind: 'binding', attr: 'attribute', action: 'action', prose: 'prose' };
const LINES = [
{ toks: [['btn', 'tag'], ['Decrement', 'label'], ['ghost', 'mod'], ['>count--', 'action']],
verdict: 'STRUCTURE: btn is a control, so every word is a token' },
{ toks: [['p', 'tag'], ['Press the buttons to change the value.', 'prose']],
verdict: 'PROSE: p is a text tag, so the rest of the line is copy, verbatim' },
{ toks: [['p', 'tag'], ['Set x=1 to enable', 'prose']],
verdict: 'PROSE: there is an "=", but `x` is not an attribute p accepts' },
{ toks: [['text', 'tag'], ['{title}', 'bind'], ['strike={done}', 'attr']],
verdict: 'STRUCTURE: text tag, but `strike=` is an attribute it registers' },
{ toks: [['btn', 'tag'], ['Add', 'label'], ['primary', 'mod'], ['disabled={!draft.trim()}', 'attr'], ['busy="Adding…"', 'attr']],
verdict: 'STRUCTURE: label, modifier, two attributes' },
{ toks: [['head', 'tag'], ['Tasks — {tasks.open.count} open', 'prose']],
verdict: 'PROSE with an interpolated binding: head is a text tag' },
{ toks: [['form', 'tag'], ['>tasks.add{title:draft}; draft=""', 'action']],
verdict: 'STRUCTURE: `>` swallows the rest of the line, so actions lex in one pass' }
];
let li = 0, t = 0;
function setup() { createCanvas(windowWidth, 250); textFont('monospace'); }
function windowResized() { resizeCanvas(windowWidth, 250); }
function draw() {
background(22, 24, 29);
t++;
const L = LINES[li];
const shown = min(L.toks.length, floor(t / 28));
const ts = constrain(width / 40, 11, 16);
textSize(ts); textAlign(LEFT, CENTER);
let x = 20, y = 70;
for (let i = 0; i < L.toks.length; i++) {
const [s, k] = L.toks[i];
const w = textWidth(s) + 14;
if (x + w > width - 16) { x = 20; y += 58; }
const on = i < shown;
noStroke();
fill(C[k][0], C[k][1], C[k][2], on ? 40 : 0);
rect(x, y - ts, w, ts * 2, 5);
fill(on ? color(C[k][0], C[k][1], C[k][2]) : color(90));
text(s, x + 7, y);
if (on) { textSize(9); fill(C[k][0], C[k][1], C[k][2]); text(NAMES[k], x + 7, y + ts + 9); textSize(ts); }
if (i === shown - 1 && shown < L.toks.length) { fill(255, 209, 102); triangle(x + w / 2 - 5, y - ts - 12, x + w / 2 + 5, y - ts - 12, x + w / 2, y - ts - 4); }
x += w + 8;
}
if (shown >= L.toks.length) {
fill(235); textSize(constrain(width / 52, 10, 13)); textAlign(LEFT, TOP);
text(L.verdict, 20, height - 64, width - 40);
}
fill(110); textSize(11); textAlign(LEFT, TOP);
text('line ' + (li + 1) + ' of ' + LINES.length + ' - click for the next line', 20, 14);
if (t > L.toks.length * 28 + 170) next();
}
function next() { li = (li + 1) % LINES.length; t = 0; }
function mousePressed() { next(); }Two more lexer decisions I like:
- Actions terminate the line by construction. Everything after
>is one token, so actions are lexable in a single pass with no lookahead. - URLs are an allow-list. A route is
/pathat token start, so$24/mostays one word, and absolute URLs are limited tohttp://andhttps://. As the comment says, a general "scheme followed by://" rule would makejavascript:anddata:lexable as request targets. The allow-list is the security boundary.
Every error in one pass
This is the invariant I'd defend hardest, and the project's own CLAUDE.md says it plainly: "Each repair-loop round is a full LLM generation; single-error reporting is a product defect, not a style preference."
A compiler that stops at the first error is fine for a human, who fixes one thing and presses save. For a model it's a disaster: each round trip is a full generation, billed at output rates. So the parser recovers from everything:
- An unknown tag is reported, with a suggestion, and then parsed as a container anyway, so errors in its children still surface.
- A missing attribute value is reported and parsing continues.
- Bad indentation is reported once per run, not once per line. Before that fix, one mis-indented
sectionproduced 15 strayGUML0011s, and a repair loop handed sixteen errors will edit sixteen lines, fifteen of which were correct.
Here's the playground's deliberately broken sample, three mistakes from three different passes:

And the same file from the CLI, exactly as guml check printed it on my machine:
error[GUML0030]: unknown tag `crad`
--> broken.guml:4:1
|
4 | crad sm center
| ^^^^
= help: did you mean `card`?
= suggestion: card
error[GUML0033]: binding refers to `kount`, which is not declared
--> broken.guml:6:3
|
6 | metric {kount}
| ^^^^^^^^^^^^^^
= help: did you mean `count`?
= suggestion: count
error[GUML0050]: `btn` has no accessible name
--> broken.guml:7:3
|
7 | btn quiet >count++
| ^^^^^^^^^^^^^^^^^^
= help: give the control a text label, or add `aria="…"` describing what it does
= suggestion: btn aria="…"
(Look closely at the carets under metric {kount}. They'll matter later.)
Notice the third error. A button with no accessible name is a compile error, not a lint warning. The explain text says why: "an inaccessible interface is a broken interface, and the compiler is the last place able to notice."
Diagnostic codes are an API
There are 49 codes, GUML0001 to GUML0104, grouped in decades: lexical, layout, syntax, resolution, semantics, accessibility, types, structure, domains, escape hatches. They're append-only: a test enforces it, and the Rust enum is #[non_exhaustive]. The reason is that a repair loop keys on them, so GUML0033 has to mean the same thing forever. 0060 is deliberately unallocated, because "shipping a code with no emit site is how a diagnostic surface starts lying about what it can detect."
Every diagnostic is machine-actionable JSON:
{"code":"unknown_tag","id":"GUML0030","severity":"error","message":"unknown tag `crad`",
"span":{"start":27,"end":31,"line":4,"col":1},"help":"did you mean `card`?","suggestion":"card"}
suggestion is present only when the fix is unambiguous, which is what lets a machine apply it without asking anyone.
Did you mean card?
How does the compiler know crad means card? The obvious answer is edit distance, and the obvious answer is wrong. Plain Levenshtein scores crad → card as 2 (two substitutions), which is too far for a four-letter word. Typos are very often swapped neighbours, so GUML uses Optimal String Alignment, which adds one more case to the recurrence:
Watch the two tables fill, and see where the transposition (highlighted) saves an edit:
const PAIRS = [['crad', 'card'], ['kount', 'count'], ['primry', 'primary'], ['tabel', 'table']];
let pi = 0, t = 0, A, B, lev, osa, order;
function setup() { createCanvas(windowWidth, 330); textFont('monospace'); build(); }
function windowResized() { resizeCanvas(windowWidth, 330); }
function build() {
[A, B] = PAIRS[pi];
lev = table(false); osa = table(true);
order = [];
for (let i = 1; i <= A.length; i++) for (let j = 1; j <= B.length; j++) order.push([i, j]);
t = 0;
}
function table(transpose) {
const d = [];
for (let i = 0; i <= A.length; i++) { d.push([]); for (let j = 0; j <= B.length; j++) d[i].push({ v: 0, tr: false }); }
for (let i = 0; i <= A.length; i++) d[i][0].v = i;
for (let j = 0; j <= B.length; j++) d[0][j].v = j;
for (let i = 1; i <= A.length; i++) for (let j = 1; j <= B.length; j++) {
const cost = A[i - 1] === B[j - 1] ? 0 : 1;
let v = min(d[i - 1][j].v + 1, d[i][j - 1].v + 1, d[i - 1][j - 1].v + cost), tr = false;
if (transpose && i > 1 && j > 1 && A[i - 1] === B[j - 2] && A[i - 2] === B[j - 1] && d[i - 2][j - 2].v + 1 < v) { v = d[i - 2][j - 2].v + 1; tr = true; }
d[i][j] = { v, tr };
}
return d;
}
function grid(d, x0, y0, cs, title, filled) {
noStroke(); fill(200); textSize(12); textAlign(LEFT, BOTTOM);
text(title, x0, y0 - 26);
textAlign(CENTER, CENTER); textSize(min(13, cs * 0.45));
for (let j = 0; j < B.length; j++) { fill(150); text(B[j], x0 + (j + 2) * cs - cs / 2, y0 - 10); }
for (let i = 0; i < A.length; i++) { fill(150); text(A[i], x0 - 10, y0 + (i + 2) * cs - cs / 2); }
for (let i = 0; i <= A.length; i++) for (let j = 0; j <= B.length; j++) {
const x = x0 + j * cs, y = y0 + i * cs;
const k = (i - 1) * B.length + (j - 1);
const show = i === 0 || j === 0 || k < filled;
const last = i === A.length && j === B.length;
fill(d[i][j].tr && show ? color(255, 209, 102, 70) : last && show ? color(247, 59, 32, 90) : color(34, 38, 46));
rect(x + 1, y + 1, cs - 2, cs - 2, 3);
if (show) { fill(d[i][j].tr ? color(255, 209, 102) : color(220)); text(d[i][j].v, x + cs / 2, y + cs / 2); }
}
}
function draw() {
background(22, 24, 29);
t++;
const filled = min(order.length, floor(t / 6));
const cols = B.length + 1, rows = A.length + 1;
const cs = min(34, (width / 2 - 50) / cols, (height - 110) / rows);
grid(lev, 36, 64, cs, 'Levenshtein', filled);
grid(osa, width / 2 + 24, 64, cs, 'Optimal String Alignment', filled);
noStroke(); textAlign(LEFT, TOP); textSize(constrain(width / 50, 10, 13));
fill(235);
text(A + ' -> ' + B, 16, 10);
if (filled >= order.length) {
const a = lev[A.length][B.length].v, b = osa[A.length][B.length].v;
fill(235);
text('distance ' + a + ' vs ' + b + (a !== b ? ': swapping two adjacent letters is one edit, not two' : ': no transposition, both agree'), 16, height - 30, width - 32);
}
fill(110); textSize(10); textAlign(RIGHT, TOP); text('click for another typo', width - 16, 12);
if (t > order.length * 6 + 200) { pi = (pi + 1) % PAIRS.length; build(); }
}
function mousePressed() { pi = (pi + 1) % PAIRS.length; build(); }The threshold is at most 1 edit for names of four characters or fewer, and at most 2 above that. But edit distance can't help with the most common mistake a model makes, which is writing HTML out of habit: button → btn is three edits. So before any distance is computed, the registry checks a table of 39 HTML habits (div→col, span→text, button→btn, hr→divider, ul→menu, dialog→modal…), and each entry carries a note explaining why GUML doesn't have that element.
Modifiers are stricter still: exactly one edit, on a word of at least four characters. That rule exists because of a real data-loss bug, where btn Click me was "fixed" to btn Click md.
Repair without a model
Models don't return clean files. They return Here you go:, a code fence, the code, another fence, and Let me know if you want changes!. guml repair strips that packaging and applies every unambiguous fix, with no model call, ever:
| Layer | Fixes | Cost |
|---|---|---|
sanitize |
code fences, markdown rules, trailing commentary | free |
format |
indentation, tabs, spacing | free |
fix |
every unambiguous diagnostic suggestion, up to 3 rounds | free |
sequenceDiagram
participant M as LLM
participant R as guml repair
participant C as guml check
M->>R: "Here you go:" + fenced GUML + commentary
R->>R: sanitize (unwrap fence, drop trailing prose)
R->>R: format (indent, tabs)
loop up to 3 rounds
R->>C: check
C-->>R: diagnostics with suggestions
R->>R: apply unambiguous fixes, right to left
end
R->>C: final check
alt compiles
C-->>M: done, zero extra generations
else still failing
C-->>M: only the hard errors, with codes and spans
endTwo safety rules make this trustworthy. First, a layer that would increase the error count is discarded rather than kept: "deterministic" is not the same as "always an improvement". Second, a suggestion containing … is a template for a human (btn aria="…"), not a replacement, so it's shown and never applied. Splicing it in would put a literal ellipsis in the accessible name, which is worse than the original error.
I fed it the broken sample wrapped in a fence and some chat filler. The JSON report:
{"applied":["GUML0030"],"changed":true,"errorsBefore":7,"errorsAfter":2,"rounds":2,
"sanitize":{"fence":true,"rules":0,"trailing":0}}
Seven errors down to two with zero tokens spent. The fence and the commentary are gone and crad became card. What's left are the two that genuinely need a decision: is kount a typo or a missing state, and what should this button say?
What 24 lines become
Here's the part the model never writes. These excerpts are real output from guml build fixtures/b.guml, 253 lines of React and TypeScript in total.
A retry helper that knows what's safe to retry. It uses exponential backoff (300 ms, then 600 ms), only for idempotent methods, only on a network failure or a 5xx, and never after an abort:
async function retrying(url: string, init: RequestInit = {}, tries = 3): Promise<Response> {
const idempotent = !init.method || ["GET", "HEAD", "PUT", "DELETE"].includes(init.method);
let wait = 300;
for (let attempt = 1; ; attempt++) {
const last = attempt >= tries || !idempotent;
try {
const res = await fetch(url, init);
if (res.ok || res.status < 500 || last) return res;
} catch (err) {
if (last || (err instanceof Error && err.name === "AbortError")) throw err;
}
await new Promise((resolve) => setTimeout(resolve, wait));
wait *= 2;
}
}
A cache with stale-while-revalidate and in-flight deduplication. It caches GETs only, treats entries as stale after 30 seconds (a constant, deliberately "not a knob"), serves the last good data if a refresh fails, and lets mutations invalidate by URL prefix, so PATCH /api/tasks/{id} invalidates /api/tasks:
async function cached<T>(url: string, init: RequestInit = {}): Promise<T> {
const method = init.method ?? "GET";
const key = `${method} ${url}`;
if (method !== "GET") return read<T>(url, init);
const hit = GUML_CACHE.get(key);
const fresh = hit && Date.now() - hit.at < GUML_STALE_MS;
if (hit && !fresh) {
void refresh<T>(key, url, init).catch(() => {});
}
if (hit) return hit.data as T;
const pending = GUML_INFLIGHT.get(key);
if (pending) return pending as Promise<T>;
return refresh<T>(key, url, init);
}
A list fetch that can be cancelled, with an alive guard so a late response never writes into an unmounted component:
const tasksList = useCallback(() => {
const controller = new AbortController();
let alive = true;
setTasksLoading(true);
setTasksError(null);
cached<Task[]>("/api/tasks", { signal: controller.signal })
.then((rows) => { if (alive) setTasks(rows); })
.catch((err: unknown) => {
if (!alive || (err instanceof Error && err.name === "AbortError")) return;
setTasksError(err instanceof Error ? err.message : "Unknown error");
})
.finally(() => { if (alive) setTasksLoading(false); });
return () => { alive = false; controller.abort(); };
}, []);
And the optimistic update with rollback, all from the word optimistic:prepend:
const tasksAdd = useCallback(
async (body: Partial<Task>) => {
const snapshot = tasks;
setTasksSaving(true);
const optimistic = { id: `tmp-${Date.now()}`, ...body } as Task;
setTasks((prev) => [optimistic, ...prev]);
try {
const res = await retrying("/api/tasks", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify(body),
});
if (!res.ok) throw new Error(`Request failed: ${res.status}`);
invalidate("/api/tasks");
const created = (await res.json()) as Task;
setTasks((prev) => prev.map((it) => (it === optimistic ? created : it)));
} catch (err: unknown) {
setTasks(snapshot);
setTasksError(err instanceof Error ? err.message : "Could not save");
} finally {
setTasksSaving(false);
}
},
[tasks],
);
stateDiagram-v2 [*] --> Applied: user submits, tmp row prepended Applied --> Confirmed: 2xx, tmp row replaced by server row Applied --> RolledBack: error, setTasks(snapshot) Confirmed --> [*]: invalidate /api/tasks RolledBack --> [*]: error banner shown
Here's that state machine running. Rows appear instantly with a tmp- id. When the server answers, they either become real or vanish. Click to change the failure rate:
const RATES = [0, 0.3, 0.7];
let ri = 1, rows = [], flights = [], t = 0, n = 3, sent = 0, ok = 0, back = 0, flash = 0, msg = '';
function setup() {
createCanvas(windowWidth, 340); textFont('monospace');
rows = [{ title: 'Freeze the v0.1 spec', id: 't3', tmp: false }, { title: 'Write ten Phase 0 specs', id: 't2', tmp: false }, { title: 'Count tokens', id: 't1', tmp: false }];
}
function windowResized() { resizeCanvas(windowWidth, 340); }
function add() {
n++;
const snapshot = rows.slice();
const row = { title: 'Task ' + n, id: 'tmp-' + n, tmp: true };
rows = [row, ...rows].slice(0, 6);
flights.push({ row, snapshot, p: 0, fail: random() < RATES[ri] });
sent++;
}
function draw() {
background(22, 24, 29);
t++;
if (t % 150 === 1) add();
const lx = 16, lw = min(300, width * 0.48), sx = width - 70;
noStroke(); textAlign(LEFT, CENTER); textSize(12);
rows.forEach((r, i) => {
const y = 64 + i * 40;
fill(r.tmp ? color(52, 58, 70, 140) : color(44, 49, 58));
rect(lx, y, lw, 32, 5);
fill(r.tmp ? 160 : 230); text(r.title, lx + 12, y + 16);
fill(r.tmp ? color(255, 209, 102) : color(110)); textSize(9); textAlign(RIGHT, CENTER);
text(r.id, lx + lw - 10, y + 16); textSize(12); textAlign(LEFT, CENTER);
});
fill(52, 58, 70); rect(sx - 28, 110, 56, 80, 6);
fill(200); textAlign(CENTER, CENTER); textSize(10); text('POST', sx, 140); text('/api/tasks', sx, 156);
for (const f of flights) {
f.p += 0.012;
const out = f.p < 0.5, k = out ? f.p * 2 : (1 - f.p) * 2;
const x = lerp(lx + lw + 10, sx - 30, k);
if (out) fill(255, 209, 102); else fill(f.fail ? color(239, 71, 111) : color(42, 157, 143));
circle(x, 150 + (out ? -10 : 10), 10);
if (f.p >= 1 && !f.done) {
f.done = true;
if (f.fail) { rows = f.snapshot; back++; flash = 60; msg = 'HTTP 500: setTasks(snapshot), the row is gone'; }
else { f.row.tmp = false; f.row.id = 'srv-' + f.row.title.split(' ')[1]; ok++; msg = 'HTTP 201: tmp row replaced by the server row'; }
}
}
flights = flights.filter((f) => !f.done);
if (flash > 0) { flash--; fill(239, 71, 111, flash * 3); rect(0, 0, width, height); }
fill(235); textAlign(LEFT, TOP); textSize(constrain(width / 48, 10, 13));
text('sent ' + sent + ' confirmed ' + ok + ' rolled back ' + back + ' failure rate ' + round(RATES[ri] * 100) + '%', 16, 12);
fill(170); textSize(11); text(msg, 16, height - 44, width - 32);
fill(110); text('click to change the failure rate - POST is not idempotent, so it is never retried', 16, height - 22, width - 32);
}
function mousePressed() { ri = (ri + 1) % RATES.length; }Leave it at 70% for a while and you'll spot a subtle thing. When two adds overlap and the first fails, the rollback restores the first one's snapshot, so the second row disappears too, even if it later succeeds. The generated React has the same property, because snapshot is captured per call. It's on my list.
Two smaller niceties: the tasksSaving flag only exists if some busy= reads it, and it clears in finally, "or one failed mutation leaves the button saying 'Adding…' forever". And an error boundary is only emitted when the document uses an escape hatch, since generated code can't throw in ways the compiler didn't write.
Seven backends, one vocabulary
The same 24 lines through every backend:
| Backend | --backend |
Lines for b.guml | What it's for |
|---|---|---|---|
| React + TS | react |
253 | the default; ~85 lines are the inlined runtime |
| React, shared runtime | react --runtime shared |
171 + 95 | one runtime module across pages, so invalidation holds across them |
| Svelte 5 | svelte |
202 | runes |
| Web Components | wc |
267 | no framework, no build step |
| Static HTML | html |
297 | mostly the inlined stylesheet |
| JSON UI tree | json |
274 | for your own renderer; the playground preview uses it |
| A2UI | a2ui |
240 | Google's agent-UI wire format, "a2ui-shaped" |
| MCP-UI | mcp-ui |
13 | one JSON resource embedding the wc script |
The design goal is that backends share decisions rather than copy them: one element table (element_for(tag)), one design-system table (theme::active().classes(tag, mods)), one expression lowering (expr::lower_in) and one liveness answer (referenced_names, which deliberately over-approximates so codegen never drops a declaration the output still mentions).
The element table exists because of a bug. The HTML backend once had a fallback of _ => "div", so the no-JavaScript build had no landmarks. Now a test checks that every tag lowers to the same element in every backend.
The A2UI emitter is labelled honestly: it was "written from a description of the protocol rather than against its published JSON schema". Actions become declared intents, so no JavaScript travels in the payload.
Themes and capabilities
A theme is data, with a contract
There are no class strings in GUML source, so all presentation comes from a theme. A theme is JSON: rules mapping (tag, modifiers) to classes, plus a contract the compiler enforces. From crates/guml-codegen/src/theme.rs:
pub fn validate(&self) -> Result<(), ThemeError> {
if self.contract.focus_visible.trim().is_empty() {
return Err(ThemeError::WeakContract(format!(
"theme `{}` declares no `contract.focus_visible`; a control nobody can see focus on \
is unusable from a keyboard, and the compiler cannot supply one it does not know", self.name)));
}
if self.contract.min_contrast < 4.5 {
return Err(ThemeError::WeakContract(format!(
"theme `{}` declares `min_contrast` {}, below the WCAG AA floor of 4.5 for body text", ...)));
}
Ok(())
}
The compiler appends the contract's focus and disabled classes to every focusable element itself, "so a theme cannot forget it on one control". That's what lets a themeable compiler still promise accessible output. Two themes are built in: stock tailwind (the default) and shadcn, which uses shadcn/ui's own oklch tokens, so a host already running shadcn drops it in unchanged.
What will this document do?
If a model writes your UI, you want to know what it can do before you render it. guml capabilities walks the AST and tells you. Real output over the fixtures:
fixtures\b.guml — Tasks (app)
network: self
GET /api/tasks (tasks)
POST /api/tasks (tasks.add) mutating
PATCH /api/tasks/{id} (tasks.save) mutating
DELETE /api/tasks/{id} (tasks.drop) mutating
component `tabs` needs a runtime
component `list` needs a runtime
fixtures\c.guml — Landing (core)
inert: no script, no network, markup only
fixtures\d.guml — Spend (app)
script: a `js` block runs code the compiler does not check
raw: host markup the compiler does not escape
escape hatches: 3 (10.7% of lines)
It also derives a Content-Security-Policy from that manifest, so the browser enforces what the compiler found:
$ guml capabilities fixtures/b.guml --csp react
default-src 'none'; connect-src 'self'; script-src 'self'; style-src 'self'; img-src 'self' data:; form-action 'none'; frame-ancestors 'none'; base-uri 'none'
flowchart LR
DOC["GUML document"] --> MAN["capabilities manifest"]
MAN --> NET["declared origins"] --> CS["connect-src"]
MAN --> SCR{"js block, or app level?"}
SCR -- "no" --> SN["script-src 'none'"]
SCR -- "yes" --> SS["script-src 'self'"]
MAN --> GATE["--assert-inert: fail CI unless core, no script, no network"]
MAN --> BUD["--max-escapes N: an escape-hatch ratchet for CI"]What is measured, and what is not
This project is a study before it's a language, and the repo is unusually strict about labelling claims. I'll keep the same discipline here, because it's easy to overclaim in this space.
Measured: authored fixtures, a real tokenizer
These are hand-written GUML and hand-written React for the same three apps, counted with cl100k_base:
| Fixture | React + TS + Tailwind | GUML | Reduction | Ratio |
|---|---|---|---|---|
| Counter card | 368 | 64 | 82.6% | 5.8× |
| Task CRUD | 1,434 | 173 | 87.9% | 8.3× |
| Landing page | 1,648 | 376 | 77.2% | 4.4× |
| Total | 3,450 | 613 | 82.2% | 5.6× |
o200k_base agrees within a percentage point. The Task CRUD row has a story. It was first published as 175 tokens. A recount said 182, and after guml fmt removed 12 bytes of padding it's 173. The research report says so in a correction box: "the original figure was wrong by 7 tokens and no one caught it, which is exactly the failure this project accuses other people's headline claims of." The culprit was Windows Python reading stdin in the console codepage, which mangled the … character.
The obvious objection is "why not just emit JSON?", so that was measured too. Minified JSON for the same task app is 315 tokens against GUML's 173, so GUML is 45% smaller.
The content floor
Compression isn't uniform, and the reason is prose. In the landing page, 232 of GUML's 376 tokens are the human copy itself, which costs the same in any language. Structure is only 144 tokens in GUML, against about 1,416 in React. So the ratio follows a simple curve:
r(P) = \frac{S_{\text{React}} + P}{S_{\text{GUML}} + P}, \qquad r(0) = \frac{1416}{144} \approx 9.8\times, \qquad \lim_{P \to \infty} r(P) = 1Drag across it:
const SG = 144, SR = 1416;
const PMAX = 2500;
let P = 232, t = 0, M = { l: 56, r: 24, t: 54, b: 70 };
function setup() { createCanvas(windowWidth, 360); textFont('monospace'); }
function windowResized() { resizeCanvas(windowWidth, 360); }
function ratio(p) { return (SR + p) / (SG + p); }
function px(p) { return map(p, 0, PMAX, M.l, width - M.r); }
function py(r) { return map(r, 1, 10, height - M.b, M.t); }
function draw() {
background(22, 24, 29);
t++;
const over = mouseX > M.l && mouseX < width - M.r && mouseY > 0 && mouseY < height;
P = over ? map(mouseX, M.l, width - M.r, 0, PMAX) : 232 + 1100 * (1 - cos(t / 160)) / 2;
stroke(48); strokeWeight(1);
for (let r = 2; r <= 10; r += 2) line(M.l, py(r), width - M.r, py(r));
noStroke(); fill(120); textSize(10); textAlign(RIGHT, CENTER);
for (let r = 2; r <= 10; r += 2) text(r + 'x', M.l - 8, py(r));
textAlign(CENTER, TOP);
for (let p = 0; p <= PMAX; p += 500) text(p, px(p), height - M.b + 8);
text('prose tokens in the page (identical in GUML and React)', (M.l + width - M.r) / 2, height - M.b + 24);
noFill(); stroke(247, 59, 32); strokeWeight(2);
beginShape();
for (let p = 0; p <= PMAX; p += 20) vertex(px(p), py(ratio(p)));
endShape();
noStroke(); fill(200); circle(px(232), py(ratio(232)), 8);
textAlign(LEFT, BOTTOM); textSize(10);
text('fixture C (landing): 1,648 / 376 = 4.4x', px(232) + 8, py(ratio(232)) - 4);
const r = ratio(P);
fill(255, 209, 102); circle(px(P), py(r), 11);
stroke(255, 209, 102, 90); strokeWeight(1); line(px(P), py(r), px(P), height - M.b);
noStroke(); textAlign(LEFT, TOP); fill(235); textSize(constrain(width / 46, 10, 14));
text('prose ' + round(P) + ' -> React ' + round(SR + P) + ' vs GUML ' + round(SG + P) + ' tokens = ' + nf(r, 1, 2) + 'x', M.l, 12);
fill(120); textSize(11);
text(over ? 'move across the chart to add prose' : 'hover the chart to drive it yourself', M.l, 32);
}Structure-heavy screens approach 8–10×, and content-heavy ones settle toward 2–3×. In the report's words: "Any benchmark that reports a single average number is misleading." The benchmark harness takes that literally and has no code path that prints an overall average.
Measured: GUML against its own compiler output
The README's headline 7.8× is a different measurement, so it gets its own label. scripts/measure.mjs compiles every fixture and divides the size of the emitted React by the size of the GUML. That needs no second author, because nobody wrote the right-hand side. It uses a 3.6 characters-per-token estimate, not a tokenizer. My run:
fixture guml emitted ratio vs hand-written
----------------------------------------------------------
a.guml 67 476 7.1× 5.8×
b.guml 183 2907 15.9× 8.8×
c.guml 420 1823 4.3× 4.2×
d.guml 203 2597 12.8× —
e.guml 604 4434 7.3× —
invoices.guml 211 3087 14.6× —
portfolio.guml 1039 5866 5.6× —
----------------------------------------------------------
total 2727 21190 7.8×
The script prints its own caveat: the ratio rises when the compiler generates more code. b.guml reads 15.9× here because the compiler now emits a cache layer a human wouldn't write, so 15.9× and the 8.3× above must never be quoted together.
Measured: the cost of a change
First drafts aren't where the time goes. bench/guml-bench/edit-locality.mjs scripts eight real modifications and compares the smallest unified diff in each language. That's the fair baseline, diff-based editing, not regeneration. Hover the bars:
const data = [
{ edit: 'add-empty-state', guml: 7, react: 95 },
{ edit: 'add-filter', guml: 24, react: 259 },
{ edit: 'add-mutation', guml: 30, react: 266 },
{ edit: 'add-column', guml: 4, react: 22 },
{ edit: 'rename-field', guml: 16, react: 36 },
{ edit: 'restyle-control', guml: 30, react: 59 },
{ edit: 'add-prose', guml: 32, react: 57 }
].map((d) => ({ ...d, ratio: d.react / d.guml }));
const dark = window.matchMedia && window.matchMedia('(prefers-color-scheme: dark)').matches;
const bar = dark ? '#ff6a4d' : '#f73b20';
const W = Math.max(320, el.clientWidth || 640), rowH = 34;
const m = { t: 28, r: 64, b: 34, l: 128 };
const H = m.t + m.b + data.length * rowH;
const x = d3.scaleLinear().domain([0, 14]).range([m.l, W - m.r]);
const y = d3.scaleBand().domain(data.map((d) => d.edit)).range([m.t, H - m.b]).padding(0.32);
const svg = d3.select(el).append('svg').attr('viewBox', `0 0 ${W} ${H}`).attr('width', '100%')
.attr('role', 'img').attr('aria-label', 'Edit-locality: React diff tokens divided by GUML diff tokens, seven scripted edits, median 5.5x')
.style('font', '12px ui-monospace, monospace').style('color', 'currentColor');
svg.append('g').attr('transform', `translate(0,${H - m.b})`)
.call(d3.axisBottom(x).ticks(7).tickFormat((v) => v + 'x').tickSize(-(H - m.t - m.b)))
.call((g) => g.selectAll('line').attr('stroke', 'currentColor').attr('stroke-opacity', 0.12))
.call((g) => g.select('.domain').remove())
.call((g) => g.selectAll('text').attr('fill', 'currentColor').attr('fill-opacity', 0.65));
const med = 5.5;
svg.append('line').attr('x1', x(med)).attr('x2', x(med)).attr('y1', m.t - 8).attr('y2', H - m.b)
.attr('stroke', 'currentColor').attr('stroke-opacity', 0.55).attr('stroke-dasharray', '3 3');
svg.append('text').attr('x', x(med) + 4).attr('y', m.t - 12).attr('fill', 'currentColor').attr('fill-opacity', 0.75).text('median 5.5x');
svg.append('g').selectAll('text').data(data).join('text')
.attr('x', m.l - 10).attr('y', (d) => y(d.edit) + y.bandwidth() / 2).attr('dy', '0.35em')
.attr('text-anchor', 'end').attr('fill', 'currentColor').text((d) => d.edit);
const tip = d3.select(el).append('div').style('position', 'absolute').style('pointer-events', 'none')
.style('padding', '6px 8px').style('border-radius', '6px').style('font', '12px ui-monospace, monospace')
.style('background', dark ? '#1f1f1f' : '#fff').style('color', dark ? '#eee' : '#222')
.style('box-shadow', '0 2px 8px rgba(0,0,0,.25)').style('opacity', 0);
d3.select(el).style('position', 'relative');
const rows = svg.append('g').selectAll('g').data(data).join('g');
rows.append('rect').attr('x', m.l).attr('y', (d) => y(d.edit)).attr('height', y.bandwidth())
.attr('width', (d) => x(d.ratio) - m.l).attr('rx', 4).attr('fill', bar);
rows.append('text').attr('x', (d) => x(d.ratio) + 6).attr('y', (d) => y(d.edit) + y.bandwidth() / 2).attr('dy', '0.35em')
.attr('fill', 'currentColor').attr('fill-opacity', 0.8).text((d) => d.ratio.toFixed(1) + 'x');
rows.append('rect').attr('x', 0).attr('width', W).attr('y', (d) => y(d.edit) - 4).attr('height', y.bandwidth() + 8)
.attr('fill', 'transparent')
.on('mousemove', (e, d) => {
const [px, py] = d3.pointer(e, el);
tip.style('opacity', 1).style('left', Math.min(px + 12, W - 190) + 'px').style('top', py - 40 + 'px')
.html(`<b>${d.edit}</b><br>GUML diff ${d.guml} tok · React diff ${d.react} tok<br>${d.ratio.toFixed(2)}x`);
})
.on('mouseleave', () => tip.style('opacity', 0));The median over the seven comparable edits is 5.5×. The eighth edit, add a loading state, isn't in the chart because it costs zero tokens in GUML (a resource's loading state is generated from its declaration) against 67 in React. And again the caveat that ships with the script: it's 8 scripted modifications, where the report specifies 50 tasks × 3. It's a seed set, not a publishable sample.
Not measured
Whether a model can produce correct GUML, and whether that's better or worse than producing React. That's the actual research question, and it's still open. More on it below.
A real model tried it
There's one exploratory run with a real model, and it's reported because a negative result is worth reporting. Six small apps were generated by Llama 3.1 8B Instruct on build.nvidia.com, with the shipping prompt (about 3,073 estimated tokens), temperature 0.2, one generation each:
| App | Parses | Requirements met | First error |
|---|---|---|---|
| todo | yes | 8 / 8 | — |
| expenses | no | 6 / 6 | GUML0030 unknown tag option |
| bmi | no | 5 / 6 | GUML0020 trailing prose |
| dashboard | no | 5 / 6 | GUML0033 undeclared order |
| signup | no | 5 / 6 | GUML0030 option |
| tip | no | 3 / 5 | GUML0030 unknown tag --- |
1 of 6 parsed, but 32 of 37 functional requirements were met. The findings file says it best: "The gap between those two numbers is the finding." The model understood the apps: the right resources, mutations, aggregates, enumerated filters and disabled bindings. It failed on surface rules.
Then the interesting part. Two of those failures were the compiler refusing information it already had. A select with option children was rejected even though codegen already accepted that spelling, and a field's accessible name was sitting in the binding next to it. After fixing both, the same six generations went from 1 of 6 compiling to 3 of 6, with no regeneration: identical model output, and a compiler that stopped saying no.
The two that still fail are GUML0023, the expression language: **, and orders.where(status=new).count. That's what a model reaches for and the parser doesn't cover yet. One more humbling number: of nine model-driven repair attempts, seven did no better than the free layers, and two made things worse. A remarkable bonus is that todo.guml independently reproduced the shape of the hand-written task fixture from a prompt that never mentioned it.
The playground and the chat
The docs site is a Next.js app, and everything interactive on it runs the real compiler as WebAssembly in your tab. The playground shows real diagnostics, a format button, and an "emitted React" pane that is byte-for-byte what guml build writes:

The preview is rendered from the compiler's own JSON UI tree, never by eval, so it can't drift from the code.
The /chat page closes the loop: describe an interface, a hosted model answers in GUML, and the compiler renders it in the tab. The free repair layers run on the answer, and the page tells you what they changed. The prompt-tax meter is honest too, at about 3,713 estimated tokens, and the system prompt is generated from the same file the Phase 0 experiment uses: "A product that prompts differently from the experiment makes the experiment decorative."

For agents there's an MCP server, guml mcp, with five tools: guml_registry, guml_spec, guml_check, guml_repair and guml_compile. A prompt can carry a vocabulary, but it can't tell a model whether what it just wrote is correct. That needs the compiler:
→ guml_check(source) DOES NOT COMPILE — 1 error(s). Fix these and check again.
error [GUML0030] line 4: unknown tag `crad`
help: did you mean `card`?
replace with: card
Proof: what I ran for this post
I didn't want to restate README numbers, so I re-ran what I could on a clean checkout of main (4418bc5):
| Claim | How I checked | Result |
|---|---|---|
| 561 Rust tests | cargo test --workspace |
561 passed, 0 failed, 1 ignored |
| 49 diagnostic codes | counted from guml explain |
49 |
| 61 conformance cases | counted the cases in spec/tests/*.txt |
15 + 30 + 7 + 9 = 61 |
| 7.8× vs emitted React | node scripts/measure.mjs |
7.8×, output above |
| Edit locality | node bench/guml-bench/edit-locality.mjs |
median 5.50× over 7 edits |
| One-pass diagnostics | guml check broken.guml |
3 errors from 3 passes, output above |
| Free repair | guml repair --format json |
7 → 2 errors, no model call |
| Playground, chat, emitted React | local docs site + headless Chrome | the screenshots in this post |
Not re-run by me: the 1,000,000-document fuzz run (the one ignored test), the mutation study (1,031 definitionally invalid mutants, 95.5% detected), and anything needing an API key.
What I found while writing this
Same rule as every project post: if I find something while writing, it goes in. GUML's own invariant is "a quietly wrong compiler destroys the reliability claim the whole project rests on", so these matter more here than they would elsewhere.
1. "Apply 3 fixes" deletes a line
Click apply 3 fixes on the playground's broken sample and you get this:

metric {kount} became the bare word count, which is now an unknown tag. Remember those carets? GUML0033's span covers the whole element, but its suggestion is just the identifier. The Rust applier already knows about this failure mode and refuses it. From crates/guml-compiler/src/fix.rs:
/// A suggestion replaces its span, so a bare word attached to a whole-line span replaces the
/// line with that word. [...] the failure mode is silent data loss, so the applier checks
/// rather than trusts: if the span covers whitespace, the replacement has to as well.
fn spans_one_token(span_text: &str, replacement: &str) -> bool {
!span_text.contains(char::is_whitespace) || replacement.contains(char::is_whitespace)
}
The TypeScript twin in @guml/core, which the playground and the chat page both call, doesn't have that guard:
export function applySuggestion(source: string, d: Diagnostic): string {
if (!d.suggestion || !isApplicable(d)) return source;
return source.slice(0, d.span.start) + d.suggestion + source.slice(d.span.end);
}
Also, the button counts every diagnostic with a suggestion, including the aria="…" template it then correctly skips, so it says "3 fixes" and applies 2. The fix is small: port spans_one_token to TypeScript, count only applicable suggestions, and ideally narrow GUML0033's span to the identifier. The lesson is bigger: two implementations of one policy drift, and the test that would catch it is a parity test over the same inputs.
2. Web Components: Delete changes the filter
In the wc output for the task app, top-level actions and row actions are numbered independently, both starting at 0. So the row's Delete button carries data-g-act="1", which is exactly the case "click:1" that sets the filter to all. The row checkbox carries data-g-act="0", which only has a submit:0 case. Clicking Delete changes the filter, and ticking a task does nothing. The existing test compares sets of action indices, and the collision makes those sets equal, so it passes. MCP-UI embeds the same script, so it inherits the bug.
3. Svelte: new tasks are posted empty
The Svelte backend emits async function tasksAdd(item, body = {}) for every mutation, but the form calls tasksAdd({ title: draft }). The object lands in item, so the POST body is {}.
4. raw html passes the "inert" gate
core is described as "safe to render from an untrusted agent", and --assert-inert as the gate for that. But this document passes both --core and --assert-inert, and the HTML backend emits the handler verbatim:
page P
raw html
<img src=x onerror=alert(1)>
p Hi
The parser rejects js at the core level but not raw, and is_inert() ignores the raw_markup flag the manifest already computes. The derived CSP would block the inline handler, but only if the host actually sends it. I'd fix this one first.
5. Docs drift
- The README says shadcn/ui is the default theme; the code defaults to stock Tailwind.
- The committed wasm is 861 KB, not the 787 KB the README and playground quote.
- The docs home page says "396 tests" and the architecture page "552 tests"; the suite has 561.
The repo already has check-readme-claims.mjs to stop exactly this, and these are the claims it doesn't cover yet.
None of this changes the argument, and all of it was found by running the thing rather than reading it, which is the project's own philosophy. React is the backend that's exercised by typecheck and render tests, and it's the one that held up. The others need the same harness pointed at them.
The question still open
Everything above is the instrument. The research report is blunt about what the language itself contributes: "the language is not novel. The measurement, the crossover characterization, and the convention-as-correctness-guarantee framing are."
The tension is real. A DSL cuts output tokens and narrows what a model can get wrong, but it also moves the model off-distribution. GUML has zero training data by construction, and the low-resource-language literature consistently reports accuracy drops. Against that, Anka (arXiv:2512.23214) reports a constrained DSL beating Python by 40 points on multi-step tasks, in one narrow domain. Nobody has characterised where the crossover is.
So the report pre-registers six hypotheses: validity, correctness ("the crux hypothesis. If H2 fails, the project is a cost optimization, not a research contribution"), hallucination, consistency, the accessibility floor, and a capability threshold. The last one predicts GUML helps small models most and may invert on frontier models.
And it defines a two-week kill-or-continue experiment, Phase 0:
flowchart TD
T["10 tasks: structure-heavy, content-heavy, mixed"] --> G["90 generations: 3 models × 0 or 3 examples × GUML, plus React"]
G --> C1{"≥ 80% of Sonnet 5 generations at 3 examples parse?"}
C1 -- "no" --> K["publish the negative result"]
C1 -- "yes" --> C2{"median ≥ 3× fewer output tokens on structure-heavy tasks?"}
C2 -- "no" --> K
C2 -- "yes" --> C3{"semantic correctness not worse than React? (blind human grader)"}
C3 -- "no" --> L["tokens win, correctness loses: the low-resource penalty"]
C3 -- "yes" --> GO["continue: build GUML-Bench, 150 tasks"]Status, verbatim from the repo: "not yet run. The gate is undecided." Everything the experiment needs is built and runs in CI: the ten tasks, ten React references that typecheck under --strict, prompt assembly, scoring, and a blind scoresheet. An attempt with a valid key returned Your credit balance is too low, so zero of the 90 generations have run. What's missing is the answer, not the instrument.
What I learned
- Diagnostics are the product. In a generation loop, an error message is an API: a code, a span and a replacement. "Every error in one pass" saves real money, and "append-only codes" is a compatibility promise like any other.
- Moving work into the compiler moves bugs into the compiler. That's a good trade, because a compiler bug is fixed once for every generation. But it raises the bar: a silently wrong backend is worse than a loud model error.
- Two implementations of one rule will drift. Rust and TypeScript fix appliers, seven backends, a README and a docs site. Every drift I found was between two copies of the same decision.
- Compression has a floor, and it's prose. You can't compress the words a human has to read. Any single average ratio hides that.
- The honest number is the interesting one. 1 of 6 parsing, but 32 of 37 requirements met, taught me more than any compression figure, because it says exactly where the compiler, not the model, was the problem.
- Measure the thing that's fair, not the thing that flatters. Diff-based edits instead of regeneration, the compiler's own output instead of a second author, and a correction box when a number was wrong.
What's next
- Fix the five findings above, starting with the
rawinert gate, then portspans_one_tokento@guml/corewith a parity test. - Point the typecheck and render harness at Svelte and Web Components, not just React. The
wccollision would have been one click away from being caught. - Grow the expression language where models actually reach:
**, ternaries, and.where(...)aggregates (GUML0023). - Playwright, axe and Lighthouse over the emitted output. Today it's proven to typecheck and server-render, not proven to work when clicked. That's the largest gap in what's verified.
- Run Phase 0. 90 generations, three models, a blind grader. Then publish whichever answer comes back.
Try it
git clone https://github.com/ravikisha/guml && cd guml
cargo test --workspace
cargo run -q -p guml-cli -- build fixtures/b.guml # React + TS
cargo run -q -p guml-cli -- build fixtures/b.guml --backend svelte
cargo run -q -p guml-cli -- check fixtures/invoices.guml --format json
cargo run -q -p guml-cli -- capabilities fixtures
cargo run -q -p guml-cli -- explain GUML0050
# what a model actually returns, repaired with no model call:
printf 'Here you go:\n\n```guml\npage P\ndiv\n span Hi\n```\n' | cargo run -q -p guml-cli -- repair
node scripts/measure.mjs # the 7.8× table
node bench/guml-bench/edit-locality.mjs # the cost of a change
Or give your agent the compiler instead of a prompt:
{ "mcpServers": { "guml": { "command": "guml", "args": ["mcp"] } } }
If you've read my posts on RelaxLang (a language) and Relaxnative (Rust from Node), GUML is where those two threads meet: a language whose user is a model, and a Rust compiler that ships to the browser as WebAssembly.
⭐ Star it on GitHub. If you have API credit and want to help answer the question this project exists for, the Phase 0 harness is waiting. And if you find a document where the compiler is quietly wrong, open an issue. Those are the bugs that matter most here.
Happy compiling! 🦀

