What must survive when intelligent components hand off work?

By Jay · Active research program

Machine-Native Systems investigates how different intelligent components communicate, coordinate, authorize actions, represent uncertainty, and establish evidence.

It asks what information must cross a component boundary without assuming conversational language is the universal internal substrate. The program combines literature synthesis with bounded experiments. Its working position is still open to failure.

01 / Motivation

An answer, permission, and an outcome are different things.

My background in HVAC and construction taught me to notice hidden state and unclear handoffs. I bring that concern to software: what a component proposes, what it is allowed to do, and what actually happened need to remain distinguishable.

This program gives those concerns a testable form. A producer supplies information; a consumer uses it to decide what comes next. Their boundary is the exchange between them. The question is whether that exchange preserves useful meaning—and whether the consumer acts on it—at an acceptable full-system cost.

02 / Current position

Explicit semantics. An open hypothesis.

Working hypothesis
Consequential state, authority, attempts, and evidence may benefit from explicit semantics at component boundaries. Representational freedom inside intelligent components remains allowed.
Observed in these experiments
A contract can preserve relevant facts while a learned consumer fails to use them. E7 and E7.1 materially complicate the claim that such a boundary is semantically sufficient.
Still unresolved
Whether the approach offers useful coherence, recovery, or component substitution once information loss and integration costs are included. Those costs can outweigh the gains and weaken the position.

Neither typed-format superiority nor universal explicit-boundary failure is established.

03 / Experiment record

Each experiment made the thesis narrower.

Start with information loss. Then separate safe behavior from useful work. Finally ask whether a learned component uses the facts it already has.

E4 E4.1 E7 E7.1

  1. E4 / Representation

    Format was not the advantage.

    Read E4 results ↗
    Test
    Compare controlled prose, JSON, and hybrid representations with the same full information, plus a compact representation, in a scripted fixture world.
    Observation
    The three full-information representations tied. Compact representation lost relevant distinctions; optimistic missing-state defaults contributed to disagreement and duplicate effects.
    Revision
    The result directs attention to preserved meaning and default behavior. It provides no evidence that typed formats are superior.
  2. E4.1 / Omission & defaults

    Safe omission and useful progress are different problems.

    Read E4.1 results ↗
    Test
    Restore omitted fields while varying optimistic and fail-closed defaults, keeping the consumer and fixtures fixed.
    Observation
    Fail-closed behavior prevented duplicate effects while sacrificing useful completion. alternatives + attempt + status recovered the required behavior.
    Revision
    That combination was a minimum sufficient subset only among the tested omissions, for this consumer, fixtures, oracle, and exact receipt model. It is not a universal minimum.
  3. E7 / Learned novelty

    Available context can still be ignored.

    Read E7 results ↗
    Test
    Compare rich shared context, explicit, hybrid, and inventory-first adaptive boundaries with two fitted statistical text-model families under novelty.
    Observation
    Rich shared context outperformed the other arms. Hybrid supplied contradictory supporting facts, but learned consumers often failed to use them. The implemented adaptive retrieval changed no final decision.
    Revision
    Economical fixed/adaptive semantic sufficiency weakened for this implementation. Information availability and use are distinct. This does not establish that explicit boundaries generally fail.
  4. E7.1 / Assessment interference

    Producer assessment can interfere. The mechanism remains unresolved.

    Read E7.1 results ↗
    Test
    Vary producer-assessment presence and training reliability while keeping supporting facts fixed. An assessment bundle includes the producer's judgment and confidence signals.
    Observation
    Masking assessments improved novelty behavior in several configurations. With supporting facts unchanged, this causally establishes assessment-bundle interference there. The preregistered pooled verdict remained D / inconclusive.
    Revision
    Universal shortcut learning and assessment as the sole cause remain unestablished. Confidence fingerprints and rendering or fitting effects remain plausible. Masking also reduced some useful completion across the full set; it is not a deployment remedy.

04 / Current frontier

Diagnosis
before redesign.

E7.2 / Proposed · Not run

Can a learned component distinguish when producer assessments should yield to contradictory supporting context—and can that failure be diagnosed cleanly before introducing context-recovery architecture?

The proposed E7.2 diagnostic would control confidence signals and assessment presence. It asks whether context use can survive fallible, present assessments without sacrificing safe useful work or collapsing the rich-context control.

No context-recovery or escape-hatch architecture has been validated. The next step is a proposed test of the failure, with its outcome still open.

Inspect the E7.2 diagnostic rationale ↗

05 / Durable questions

The questions outlast the experiments.

  1. What information must survive a boundary?

    What decision-relevant meaning must survive compression, novelty, and boundary evolution?

  2. Who owns transitions and permission?

    Can intelligence change while transition ownership, authority, and accountability remain stable?

  3. What independently establishes a result?

    What establishes an outcome beyond the producer's explanation?

  4. How should uncertainty change behavior?

    How should uncertainty and dependent evidence change abstention, escalation, and action?

  5. What coordination remains coherent under faults and heterogeneity?

    Which mechanisms preserve useful coordination across faults and different components, at acceptable full-system cost?

06 / Scope of evidence

What this does not establish.

  • Small, same-author synthetic worlds. E4/E4.1 use scripted consumers. E7/E7.1 use two bounded statistical text-model families; novel locations still use familiar vocabulary.
  • Uncalibrated probabilities and trusted simulations. Authority and receipt stores are simulated and trusted. Valid permission, receipt integrity, and deterministic reproduction do not establish semantic correctness or deployment reliability.
  • Correlated executions. Crossed executions repeat combinations of cases and configurations. They are not independent population samples.
  • No demonstrated general advantage. These results establish no open-world, production-scale, general-LLM, or lifecycle-cost advantage. E7.1 masking is not a deployment-wide remedy.

07 / Inspect the record

The research repository is canonical.

This page is an authored reading of the public research summary. Protocols, results, raw artifacts, and publication history remain in the research record.

Source reviewed .
Public summary revision: 393a2ceb7c8372e4e020c81c782af1c12ade96bb.
Accumulated research reviewed by that summary: 2b9a04cf0920b54f7b994e2e2b67b5707d5192e6.

Return to Jay's selected work