AXM / public interoperability testAll tools

Keep the browser.
Re-check the procedure.

One disposable task. Four versions of the same screen. Can your caller reuse a known procedure, stop when its meaning changes, and catch a save that only looks successful?

Open the fixtureInspect the pinned contract

For an existing Hronaut-connected agent. No new integration or shared account.

Your browser, your workspace.

Create one scratch workspace in your own Hronaut instance. The test never connects to someone else’s browser, asks for credentials or sends the receipt anywhere.

The page keeps synthetic draft records locally. Your caller owns the decision to act and the check of what was actually saved.

A small test with a visible failure.

A layout change should preserve a valid procedure. A changed action should stop it. A misleading success message should fail the result check.

Only the receipt needs to cross between systems, after you review it.

The four cases

CaseExpected resultSave attempts
BaselinePASS: requested draft persisted1
Reordered layoutPASS: semantics survived1
Changed action meaningUNKNOWN: stop before task edits0
False success messageFAIL: saved quantity is wrong1, no retry

Those are expected outcomes, not prefilled results. This is an inspectable conformance test, not a blind benchmark. A task FAIL is the correct outcome for the last case.

The copy-ready prompt

Save the prompt as text

NEXT

Now stale the browser lifetime without changing the page meaning.

The first run changed semantics while the workspace stayed live. The mirror test keeps the procedure semantically valid, reloads the same tab after a native Hronaut continuity checkpoint, and asks whether Hronaut vetoes the write until its own lifetime boundary is reconciled.

Expected: semantics still MATCH; Hronaut continuity BLOCKED; write blocked; fixture untouched; reconcile normally; fresh semantic admission; one verified save; PASS.

What we have actually run

Reference checks passed: 30 pure-rule tests and all four expected browser outcomes, including reload continuity, saved-record read-back and receipt export. The run used Playwright 1.57.0 with Chromium 143.0.7499.4. Independent Hronaut-native run: maintainer-reported on Hronaut 2.4.28 as PASS, PASS, UNKNOWN with 0 saves, then FAIL after one save; reload continuity held and no repairs were reported.

Inspect the reference CI run · Read its result · Read the maintainer-reported native observation

What this result can support

The contract tests reuse admission and fresh saved-record checks across controlled changes, in one workspace. It also checks the fixture marker across a page reload. The marker is not Hronaut’s workspace generation; the caller must verify its own workspace identity.

The reference executor uses Playwright, not Hronaut. Yevhen Tienkaiev subsequently reported the same four outcomes natively in one Hronaut 2.4.28 scratch workspace/tab. That is attributed maintainer evidence, not an AXM-operated run. The v2 mirror test now targets Hronaut's lifetime guard while page semantics remain unchanged. This test does not establish reconnect behavior, dispatch-time race protection, malicious-client resistance or production safety.

Publication PR · Versioned SHA-256 manifest · Reference browser checks · Maintenance and evidence boundaries

Capsule source: f6c8a39105398e916f3ab0197258f9ce43f9cb2a