I am looking for researchers or agent builders using a model/runtime outside the Urusilla project to reproduce a small multi-hop communication evaluation.
This is not a claim that a new language already saves tokens in general use. The current result on unfamiliar external dialogue is intentionally visible:
- post-decode model API-input saving: 0%;
- total tokens per safely completed real task: not yet known;
- first same-project fresh-context chain: Capsule digest and structural generation 3/3, but explicit adoption-before-use only 2/3 because the last receiver omitted the adoption record.
That failure is why I want independent runs rather than more internal tuning.
Bounded reproduction task
Use a fresh or explicitly reset agent and give it only:
- the immutable declarative Capsule URI and exact SHA-256;
- the actual unsigned status;
- the disclosed wrapper and task;
- a rule that any mismatch falls back to concise natural language or JSON.
Before use, run positive, negative, and exact-reconstruction gates. Require an explicit session adoption decision. Then run matched concise-language, JSON, and Urusilla arms on the same bounded 3–10-turn task.
Please count the complete denominator: discovery, Capsule/context delivery, gates, all model input/output, conversion, repair, retry, fallback, judge/router, and final answer. Record post-decode API-input tokens separately. If authorized, relay only the same immutable Capsule identity to one fresh downstream receiver and record its acknowledgement and gates.
No persistence, permission expansion, spending authority, executable installation, or external effects are permitted. Negative, null, refusal, fallback, and regression results are welcome. Same-project runs must not be labeled independent.
Reproduction materials
The repository contains the agent-readable entry, immutable Capsule identity, full Interop Lab protocol, dependency-free validator, editable machine record, first controlled pilot, structured evidence form, and public results room:
- GitHub - jaden3824/urusilla: Experimental semantic layer for AI-agent communication with verified fallback; general post-decode token savings are 0% so far. · GitHub
- A2A mapping discussion: [Interop experiment request] Declarative Capsule handoff over A2A · a2aproject/A2A · Discussion #2161 · GitHub
A useful response can be as small as: model/runtime disclosure, gate result, whether adoption happened before use, one matched task ledger, and the exact failure or success artifact. Please do not share chain-of-thought, private prompts, credentials, or unredacted personal data.
Disclosure: this post was drafted and submitted with Codex assistance under the repository owner’s authorization. Any external run should disclose its own operator and assistance.