Notes from the inside.
Field notes on building the Twin: constructing trusted business context from real systems, governing what AI can rely on, choosing models, and doing the forward-deployed work that makes it useful.
You Are Not the Front Door
Why the user's chosen agent may become the primary entry point to software, and why every company will need a governed Twin between those agents and its systems.
Your AI Doesn't Need More Context. It Needs Approved Meaning.
Enterprise AI is not starved of data; it is starved of approved organisational meaning. Why the MLX Twin is a governed interpreter rather than a chatbot, a data dump, or a larger context layer.
Five Gates Between a Guess and a Governed Answer
A number in most AI workflows becomes true the moment the model says it. This is the governance machinery, five distinct gates, through which candidate business meaning becomes a usable Twin product, and why each gate answers a question the others cannot.
One Twin, Many Interfaces
A concrete map of how organisations, applications, task agents, fleet operators and evaluation runners reach the Twin today, why the naming had to stabilise before anything external is frozen, and the hard boundary that no third-party agent can connect to it directly yet.
Every AI Answer Should Come With a Receipt
A receipt is not a prose citation the model wrote about itself. It is a structured, machine-checkable output contract: product identity, execution identity, result identity, and an honest boundary marking exactly where proof ends.
A Connector Is Not Context
Connecting Xero, Slack or SharePoint gives an AI system data access, not an agreed definition of your business. Here is how MLX separates connectors, plugins and governed Twin products.
The Agent Is the Easy Part
As the agent layer becomes easier to buy and replace, value moves toward the company context no vendor can supply: reviewed definitions, bounded access, and evidence that survives the model.
Your Agent Remembers You. Who Remembers the Company?
Why persistent agent memory is useful but insufficient for organisational truth, and why a company needs governed definitions that are committed rather than merely recalled.
A2A Solves How Agents Talk. It Does Not Solve Who Is Allowed.
Why agent interoperability needs more than authentication, and how signed subject identity, capability intersection, governed products, and receipts prevent a helpful agent from becoming a confused deputy.
Context Is a Security Boundary, Not a Token Budget
Why context assembly is an authorisation decision, how untrusted evidence becomes an attack surface, and why deterministic capability and data boundaries must sit outside the model.
A Skill Is a Hypothesis Until It Passes an Evaluation
Why well-written agent skills can help, do nothing, or make performance worse, and how product-bound evaluations turn procedural guidance into evidence.
CubeSandbox Lands: A Useful Signal for Customer-Controlled Agent Execution
What CubeSandbox could change for isolated agent execution, how to validate its vendor-reported claims, and why sandboxing remains separate from governed Twin data access.
Agentic Finance Workflows on SunSystems: A Forward-Deployed Pattern
How MLX uses read-only SunSystems evidence for governed turnover, balance, journal, debtor, creditor, and management-reporting products without claiming ledger mutation.
Agentic Finance Workflows on Xero: A CFO and Project-Lead Pattern
How MLX syncs Xero evidence into the warehouse, publishes approved finance products through the Twin, and keeps Xero creation, approval, reconciliation, and sending explicitly out of scope today.
Kimi K2.6 Lands: Why Open-Weights Models Change the Finance AI Calculus
What Kimi K2.6 means for finance-model evaluation, and why governed Twin contracts make provider comparisons more useful without implying a current MLX integration.
Letting an LLM Write SQL Against Your Warehouse Safely
The two MLX warehouse paths: compiled read-only queries over an activated Twin product, and guarded model-authored SQL for explicit exploration.
What 'Customer-Controlled AI' Actually Means for Finance
A practical definition of customer-controlled AI across four planes: data, models, execution, and the governed business knowledge in your Twin.
Skills, Not Prompts: How Forward-Deployed Engineers Codify Finance Processes
How MLX separates approved business meaning, operational skills, deterministic guardrails, and evaluations so finance workflows do not depend on a clever prompt.
On-Prem AI for Finance: A Practical Path for Regulated Teams
A practical acceptance framework for customer-operated finance AI: deployment boundaries, external dependencies, governed Twin products, validation, evaluation, and operational proof.
LLM-Agnostic by Design: Why Finance AI Shouldn't Be Locked to One Vendor
How MLX separates governed Twin products from model routing, what multi-provider support does and does not mean today, and why switching models still requires evidence.
Why Your Operational Reporting Is Lying to You
Why siloed ledger, pipeline, and operations data breaks reporting, and how a governed Twin turns fragmented evidence into approved, source-backed business products.
Turn your systems into an approved Twin.
Connect scoped evidence, review the business definition, publish a versioned product, and let permitted AI query it through a governed read-only route.
Get in touch
team@mercurylabs.io
Deploy
Managed · read-only start
From
Mercury Labs · London