What must survive when intelligent components hand off work?
By Jay · Active research program
Machine-Native Systems investigates how different intelligent components communicate, coordinate, authorize actions, represent uncertainty, and establish evidence.
It asks what information must cross a component boundary without assuming conversational language is the universal internal substrate. The program combines literature synthesis with bounded experiments. Its working position is still open to failure.
01 / Motivation
An answer, permission, and an outcome are different things.
My background in HVAC and construction taught me to notice hidden state and unclear handoffs. I bring that concern to software: what a component proposes, what it is allowed to do, and what actually happened need to remain distinguishable.
This program gives those concerns a testable form. A producer supplies information; a consumer uses it to decide what comes next. Their boundary is the exchange between them. The question is whether that exchange preserves useful meaning—and whether the consumer acts on it—at an acceptable full-system cost.
02 / Current position
Explicit semantics. An open hypothesis.
- Working hypothesis
- Consequential state, authority, attempts, and evidence may benefit from explicit semantics at component boundaries. Representational freedom inside intelligent components remains allowed.
- Observed in these experiments
- A contract can preserve relevant facts while a learned consumer fails to use them. E7 and E7.1 materially complicate the claim that such a boundary is semantically sufficient.
- Still unresolved
- Whether the approach offers useful coherence, recovery, or component substitution once information loss and integration costs are included. Those costs can outweigh the gains and weaken the position.
Neither typed-format superiority nor universal explicit-boundary failure is established.
03 / Experiment record
Each experiment made the thesis narrower.
Start with information loss. Then separate safe behavior from useful work. Finally ask whether a learned component uses the facts it already has.
E4 E4.1 E7 E7.1
-
- Test
- Compare controlled prose, JSON, and hybrid representations with the same full information, plus a compact representation, in a scripted fixture world.
- Observation
- The three full-information representations tied. Compact representation lost relevant distinctions; optimistic missing-state defaults contributed to disagreement and duplicate effects.
- Revision
- The result directs attention to preserved meaning and default behavior. It provides no evidence that typed formats are superior.
-
E4.1 / Omission & defaults
Safe omission and useful progress are different problems.
Read E4.1 results ↗- Test
- Restore omitted fields while varying optimistic and fail-closed defaults, keeping the consumer and fixtures fixed.
- Observation
-
Fail-closed behavior prevented duplicate effects while sacrificing useful completion.
alternatives + attempt + statusrecovered the required behavior. - Revision
- That combination was a minimum sufficient subset only among the tested omissions, for this consumer, fixtures, oracle, and exact receipt model. It is not a universal minimum.
-
- Test
- Compare rich shared context, explicit, hybrid, and inventory-first adaptive boundaries with two fitted statistical text-model families under novelty.
- Observation
- Rich shared context outperformed the other arms. Hybrid supplied contradictory supporting facts, but learned consumers often failed to use them. The implemented adaptive retrieval changed no final decision.
- Revision
- Economical fixed/adaptive semantic sufficiency weakened for this implementation. Information availability and use are distinct. This does not establish that explicit boundaries generally fail.
-
E7.1 / Assessment interference
Producer assessment can interfere. The mechanism remains unresolved.
Read E7.1 results ↗- Test
- Vary producer-assessment presence and training reliability while keeping supporting facts fixed. An assessment bundle includes the producer's judgment and confidence signals.
- Observation
- Masking assessments improved novelty behavior in several configurations. With supporting facts unchanged, this causally establishes assessment-bundle interference there. The preregistered pooled verdict remained D / inconclusive.
- Revision
- Universal shortcut learning and assessment as the sole cause remain unestablished. Confidence fingerprints and rendering or fitting effects remain plausible. Masking also reduced some useful completion across the full set; it is not a deployment remedy.
04 / Current frontier
Diagnosis
before redesign.
E7.2 / Proposed · Not run
Can a learned component distinguish when producer assessments should yield to contradictory supporting context—and can that failure be diagnosed cleanly before introducing context-recovery architecture?
The proposed E7.2 diagnostic would control confidence signals and assessment presence. It asks whether context use can survive fallible, present assessments without sacrificing safe useful work or collapsing the rich-context control.
No context-recovery or escape-hatch architecture has been validated. The next step is a proposed test of the failure, with its outcome still open.
Inspect the E7.2 diagnostic rationale ↗05 / Durable questions
The questions outlast the experiments.
-
What information must survive a boundary?
What decision-relevant meaning must survive compression, novelty, and boundary evolution?
-
Who owns transitions and permission?
Can intelligence change while transition ownership, authority, and accountability remain stable?
-
What independently establishes a result?
What establishes an outcome beyond the producer's explanation?
-
How should uncertainty change behavior?
How should uncertainty and dependent evidence change abstention, escalation, and action?
-
What coordination remains coherent under faults and heterogeneity?
Which mechanisms preserve useful coordination across faults and different components, at acceptable full-system cost?
06 / Scope of evidence
What this does not establish.
- Small, same-author synthetic worlds. E4/E4.1 use scripted consumers. E7/E7.1 use two bounded statistical text-model families; novel locations still use familiar vocabulary.
- Uncalibrated probabilities and trusted simulations. Authority and receipt stores are simulated and trusted. Valid permission, receipt integrity, and deterministic reproduction do not establish semantic correctness or deployment reliability.
- Correlated executions. Crossed executions repeat combinations of cases and configurations. They are not independent population samples.
- No demonstrated general advantage. These results establish no open-world, production-scale, general-LLM, or lifecycle-cost advantage. E7.1 masking is not a deployment-wide remedy.
07 / Inspect the record
The research repository is canonical.
This page is an authored reading of the public research summary. Protocols, results, raw artifacts, and publication history remain in the research record.
- Machine-Native Systems · canonical repository ↗
- Reviewed public research summary ↗
- E4 · representation results ↗
- E4.1 · omission and defaults results ↗
- E7 · learned-boundary results ↗
- E7.1 · causal results and pooled verdict ↗
- Durable research questions ↗ · Hypotheses and falsification criteria ↗
- E7.1 postrun audit ↗ · E7.1 publication ancestry ↗
Source reviewed .
Public summary revision:
393a2ceb7c8372e4e020c81c782af1c12ade96bb.
Accumulated research reviewed by that summary:
2b9a04cf0920b54f7b994e2e2b67b5707d5192e6.