axm-core is the runtime between sealed shards and spoken questions. Spectra translates natural language to SQL. Forge compiles raw documents into shards. DuckDB executes cross-shard joins in memory. No inference at query time.
▶ Watch the exit demo — an ontology leaving a silo, sealedaxm-core is the hub package. It neither signs shards (that's axm-genesis) nor produces them from conversations (that's a spoke). It does one thing: turns a pile of sealed JSONL shard directories into a queryable relational database, and lets you ask it questions in English.
Natural language query engine. Parses intent, routes to a named query in the INTENT_ROUTES table, executes parameterized SQL against DuckDB. No LLM call at query time — the LLM runs once at distill time to build the index.
Document compiler. Takes PDFs, markdown, plain text, CSV, XLSX. Extracts structured claims at four tiers. Delegates signing and Merkle construction to axm-genesis via compile_generic_shard. Never reimplements the kernel.
The query substrate. Loads every shard's canonical JSONL tables into DuckDB at mount time. Cross-shard union views by table name. Cross-shard JOINs on episode_id and claim_id. All in memory. No separate database server. No schema migrations.
Spectra is deliberately not a RAG system. It doesn't embed your query, search for similar chunks, and ask an LLM to synthesize them. That path is expensive and non-deterministic.
Instead: your natural language query is classified into a named intent. The intent maps to a parameterized SQL template. DuckDB executes it. The output is a result set — structured, reproducible, auditable.
The intent classifier runs locally via Ollama at distill time. At query time the only thing that runs is SQL.
"What decisions have we made about the Merkle format?"
Regex + keyword rules classify into a named intent. INTENT_ROUTES maps each intent to a SQL template. No LLM call.
Named parameters substituted safely. Query executes against the mounted DuckDB tables. Zero injection surface.
Rows returned. Same query. Same shards. Same output. Always.
# Mount time (not query time): every shard's canonical JSONL tables — # graph/claims.jsonl, graph/entities.jsonl, ext/*.jsonl — load straight # into DuckDB. Routes below just query the mounted tables directly. INTENT_ROUTES = { "decisions": { "match": lambda q: "decision" in q or "decided" in q, "sql": """ SELECT episode_id, decision_text, rationale, shard_id FROM decisions WHERE lower(decision_text) LIKE lower(:term) ORDER BY created_at DESC """, "params": lambda q: {"term": extract_term(q)}, }, "failures": { "match": lambda q: "fail" in q or "graveyard" in q or "tried" in q, "sql": """ SELECT episode_id, problem, graveyard, solution FROM engineering WHERE graveyard IS NOT NULL ORDER BY created_at DESC """, "params": lambda q: {}, }, # ... additional routes } def route(query: str) -> tuple[str, dict]: q = query.lower() for name, route in INTENT_ROUTES.items(): if route["match"](q): return route["sql"], route["params"](q) raise ValueError(f"No route for: {query}")
Forge is the document ingestion pipeline. You point it at a document. It extracts structured claim candidates at four tiers — from lossless schema lift to LLM-assisted extraction — then delegates all signing and Merkle construction to compile_generic_shard. Forge never reimplements the kernel.
The output is structurally identical to a chat shard. Same manifest.json. Same canonical JSONL tables. Same verification path. Forge doesn't know it's compiling a field manual versus a legal statute versus a design document. It applies the same protocol to all of them.
Extract text blocks from PDF, markdown, or plain text. Segment into claim-sized units.
BLAKE3 each claim. Build Merkle tree from leaf hashes. Compute root hash.
axm-genesis signs the manifest via compile_generic_shard. Forge delegates — it never calls the signing primitives directly.
Claims written to graph/claims.jsonl. Manifest written. Shard directory sealed.
# Compile a PDF into a sealed Knowledge Shard axm-forge compile ./fm21-11-first-aid.pdf \ --signing-key keys/publisher.pem \ --shard-id fm21-11-hemorrhage-v1 # Output ✓ Parsed 312 claim blocks ✓ Merkle root: a3f9c2...d4e1 ✓ Signed ML-DSA-44 · FIPS 204 ✓ Written ~/.axm/shards/fm21-11-hemorrhage-v1/ # Verify immediately axm-verify shard ~/.axm/shards/fm21-11-hemorrhage-v1/ \ --trusted-key keys/publisher.pub # → {"status":"PASS","error_count":0,"errors":[]}
DuckDB runs in-process. There is no separate database server to start, configure, or migrate. When a shard mounts, axm-core loads its canonical JSONL tables — graph/claims.jsonl, graph/entities.jsonl, ext/*.jsonl — straight into DuckDB tables, then rebuilds the cross-shard union views.
New shards become queryable the moment they're mounted. No ingestion step. No index rebuild. Just canonical JSONL parsed into DuckDB and unioned by name across every mounted shard.
import duckdb, json conn = duckdb.connect() # Mount time: each shard's canonical JSONL tables load straight # into a per-shard DuckDB table. No external file format, no glob. def load_table(view_name, jsonl_path): rows = [json.loads(line) for line in open(jsonl_path)] conn.execute(f"CREATE TABLE {view_name} AS SELECT * FROM rows") for shard in mounted_shards: load_table(f"episodes__{shard.id}", shard.path / "ext/episodes@1.jsonl") load_table(f"engineering__{shard.id}", shard.path / "ext/engineering@1.jsonl") # Cross-shard union view: bare "episodes" spans every mounted shard conn.execute(""" CREATE VIEW episodes AS SELECT * FROM episodes__shard_a UNION ALL SELECT * FROM episodes__shard_b """) # Cross-shard join: decisions + their lineage — bare table names, # no glob in sight conn.execute(""" CREATE VIEW decisions AS SELECT d.*, l.superseded_by, l.reason FROM decisions d LEFT JOIN lineage l USING (episode_id) """) # New shards mount into their own per-shard tables; the union # views rebuild automatically. No re-ingestion required.
The refs table — mounted from each decision shard's ext/references@1.jsonl — carries pointers into other shards by shard_id and episode_id. DuckDB can resolve these at query time without a central registry.
-- One row per cross-shard link episode_id VARCHAR -- local episode target_shard_id VARCHAR -- foreign shard target_episode_id VARCHAR -- foreign episode claim_text VARCHAR -- the linked claim reference_type VARCHAR -- supports | supersedes -- contradicts | cites
-- Find all decisions and their source conversations SELECT d.decision_text, r.target_shard_id AS source_shard, r.claim_text AS supporting_claim, r.reference_type FROM decisions d JOIN refs r ON d.episode_id = r.episode_id WHERE r.reference_type = 'supersedes' ORDER BY d.created_at DESC
-- Walk the supersession chain for the Merkle format decision WITH RECURSIVE chain AS ( -- anchor: find the original decision SELECT episode_id, decision_text, superseded_by, 0 AS depth FROM decisions WHERE lower(decision_text) LIKE '%merkle%' AND superseded_by IS NULL UNION ALL -- recurse: find what superseded it SELECT d.episode_id, d.decision_text, d.superseded_by, chain.depth + 1 FROM decisions d JOIN chain ON d.superseded_by = chain.episode_id ) SELECT * FROM chain ORDER BY depth
Spectra classifies these queries and executes parameterized SQL. No LLM call at query time.
Platforms don't hold your data hostage — they hold your structure hostage. The object types, the links, the property semantics: the ontology is the part that doesn't come out when you export a CSV. The Ontology Exit takes it out.
You run three GETs against your own Palantir Foundry tenant — its published Ontology API v2, your credentials, never ours — save the JSON, and run one command. Out the other side: a genesis-sealed shard where every object type, typed property, primary key, link, and cardinality is a queryable claim, the verbatim API responses are preserved byte-for-byte, and the whole record verifies detached — with Palantir removed, with AXM removed, with everything removed except the bytes, the kernel, and one out-of-band key.
# your tenant, your token, three GETs (see ONTOLOGY_EXIT.md) curl .../api/v2/ontologies/<ont>/objectTypes > capture/objectTypes.json # one command: seal, verify detached, done axm-exit capture/ --out exit/ → verify=PASS object_types=3 claims=44
# the exited ontology is a working system, not an archive axm capture ontology-exit -p capture=./capture axm ask <shard> --key <pub> "what links to Aircraft" → rows, with a custody-verified provenance footer
Who needs a tested exit? EU and UK financial entities under DORA must hold exit strategies for critical ICT providers that are tested, reviewed annually — inspectors ask for the exit strategy with the most recent test results. Public bodies facing break-clause decisions need a demonstrated migration path, not a recommendation. Everyone else renewing a platform contract needs the credible option to leave, which only a tested exit provides. A sealed, detached-verifiable exit shard is a test result you can hand an auditor.
Evidence tier, stated plainly: reconciled against Palantir's published Ontology API v2 wire shapes and proven end-to-end against a sample in that documented shape — not yet run against an authorized live tenant. The first design partner's tenant completes that proof. No Palantir code runs anywhere in this path.
▶ watch the exit demo — silo → pull → sovereign record, with the real shard id
Ready to run it? First hour: clone → sealed exit · pipeline exit: schemas + dependency DAG · the ship of theseus ▶ (replace Foundry plank by plank) · what a Foundry exit does & doesn't cover · start here (the 8-repo map)
# Install axm-genesis first (required dep) pip install -e ./axm-genesis # Install axm-core pip install -e ./axm-core # Optional: install a spoke to generate shards pip install -e ./axm-chat
# Compile a document into a Knowledge Shard axm-forge compile ./document.pdf \ --signing-key keys/publisher.pem # Verify it immediately axm-verify shard ~/.axm/shards/document-v1/ \ --trusted-key keys/publisher.pub → {"status":"PASS","error_count":0}
# Natural language via Spectra axm-core spectra "what decisions have we made" axm-core spectra "what failed last week" axm-core spectra "what changed since january" # Raw DuckDB SQL directly axm-core query --sql "SELECT * FROM episodes LIMIT 5" # All shards. One DuckDB. In memory.
# axm-genesis — cryptographic kernel # BLAKE3 + ML-DSA-44 + axm-verify # axm-core — this package # Spectra + Forge + DuckDB runtime + Foundry Exit # axm-chat — spoke: conversation shards # axm-show — spoke: drone show telemetry # axm-embodied — spoke: robot sensor streams # All shards. Same format. Same verifier.