Notes from the inside.
Ideas, evidence and working examples from building the Twin. Read the thinking. Explore how it works.
I Built a Custom Blackwell Kernel for Qwen3.8-27B
I wanted more than 200 tokens per second from Qwen3.8-27B on one RTX PRO 6000. Getting there meant going below the serving framework.
George Pullen · 8 min read
Read the articleAll the field notes.
29 articles
How a Digital Twin Is the Key to Running a Portfolio of Fast-Growing Companies
How I run four fast-growing companies from one seat, what a digital twin actually is, and why playbooks - not the model - turned out to be the unlock.
Three Cents of Probing Says Jev Is a Small Qwen in a Very Good Suit
I sent Jev about 120 requests and roughly $0.03 of API spend. The token counts, the 32,768 context limit and the punctuation behaviour say pretrained Qwen-class LLM, customised tokenizer and parallel scoring heads. No new science, but genuinely good engineering.
What Should a Twin Learn From a Mistake?
An overdue invoice exposes a missing question. Could a Twin learn to ask it next time? A proposal for context that improves through evaluated, reversible changes.
I Built a Custom Blackwell Kernel for Qwen3.8-27B
I wanted more than 200 tokens per second from Qwen3.8-27B on one RTX PRO 6000. Getting there meant going below the serving framework.
China Is Winning the Post-Transformer Race
I am bullish on Moonshot, Alibaba and Z.ai because they are releasing architectural progress that the rest of us can actually build on.
The Token Is a Narrow Feedback Channel
A close reading of the full-bandwidth transformer: what latent feedback changes, what the 1B-parameter experiments establish, and why hidden computation cannot replace governed evidence.
RFQ1: Building a Relation-First Query Language for Project Venus
How RFQ1 represents relational queries, what its ten clauses mean, and how grammar-guided token masking can restrict a model to the language by construction.
Adapting Qwen3-VL Representations for Open-Set Continual Visual Learning
A four-stage investigation into whether Qwen's internal visual representations could support few-shot, open-set continual learning.
You Are Not the Front Door
Why the user's chosen agent may become the primary entry point to software, and why every company will need a governed Twin between those agents and its systems.
Your AI Doesn't Need More Context. It Needs Approved Meaning.
When sales, finance and delivery use the same words differently, retrieval alone cannot settle the report. Make the definition explicit, versioned and reusable.
Five Gates Between a Guess and a Governed Answer
Follow a recurring-revenue definition through five distinct controls, and see why evidence from one cannot stand in for another.
One Twin, Many Interfaces
A practical map of how enabled interfaces reach the Twin, what stays consistent across them, and why a connection never grants itself wider access.
Every AI Answer Should Come With a Receipt
A receipt is not a prose citation the model wrote about itself. It is a structured, machine-checkable output contract: product identity, execution identity, result identity, and an honest boundary marking exactly where proof ends.
A Connector Is Not Context
Connecting Xero, Slack or SharePoint gives an AI system data access, not an agreed definition of your business. Here is how MLX separates connectors, plugins and governed Twin products.
The Agent Is the Easy Part
As the agent layer becomes easier to buy and replace, value moves toward the company context no vendor can supply: reviewed definitions, bounded access, and evidence that survives the model.
Your Agent Remembers You. Who Remembers the Company?
Why persistent agent memory is useful but insufficient for organisational truth, and why a company needs governed definitions that are committed rather than merely recalled.
A2A Solves How Agents Talk. It Does Not Solve Who Is Allowed.
Why agent interoperability needs more than authentication, and how signed subject identity, capability intersection, governed products, and receipts prevent a helpful agent from becoming a confused deputy.
Context Is a Security Boundary, Not a Token Budget
Why context assembly is an authorisation decision, how untrusted evidence becomes an attack surface, and why deterministic capability and data boundaries must sit outside the model.
A Skill Is a Hypothesis Until It Passes an Evaluation
Why well-written agent skills can help, do nothing, or make performance worse, and how product-bound evaluations turn procedural guidance into evidence.
CubeSandbox Lands: A Useful Signal for Customer-Controlled Agent Execution
What CubeSandbox could change for isolated agent execution, how to validate its vendor-reported claims, and why sandboxing remains separate from governed Twin data access.
Agentic Finance Workflows on SunSystems: A Forward-Deployed Pattern
How MLX uses read-only SunSystems evidence for governed turnover, balance, journal, debtor, creditor, and management-reporting products without claiming ledger mutation.
Agentic Finance Workflows on Xero: A CFO and Project-Lead Pattern
How MLX syncs Xero evidence into the warehouse, publishes approved finance products through the Twin, and keeps Xero creation, approval, reconciliation, and sending explicitly out of scope today.
Kimi K2.6 Lands: Why Open-Weights Models Change the Finance AI Calculus
What Kimi K2.6 means for finance-model evaluation, and why governed Twin contracts make provider comparisons more useful without implying a current MLX integration.
Letting an LLM Write SQL Against Your Warehouse Safely
The two MLX warehouse paths: compiled read-only queries over an activated Twin product, and guarded model-authored SQL for explicit exploration.
What 'Customer-Controlled AI' Actually Means for Finance
Four practical questions for finance teams: where evidence lives, which model receives it, what an agent may do and who owns the definition.
Skills, Not Prompts: How Forward-Deployed Engineers Codify Finance Processes
How MLX separates approved business meaning, operational skills, deterministic guardrails, and evaluations so finance workflows do not depend on a clever prompt.
On-Prem AI for Finance: A Practical Path for Regulated Teams
How to assess a customer-operated finance AI deployment, agree responsibilities and prove one useful read-only workflow before expanding.
LLM-Agnostic by Design: Why Finance AI Shouldn't Be Locked to One Vendor
How MLX separates governed Twin products from model routing, what multi-provider support does and does not mean today, and why switching models still requires evidence.
Why Your Operational Reporting Is Lying to You
Why siloed ledger, pipeline, and operations data breaks reporting, and how a governed Twin turns fragmented evidence into approved, source-backed business products.
More notes, as we build.
Get the research briefing, or follow new articles in your own reader.
Start with the question your team keeps chasing.
Your first workflow
Show us a recurring report, an awkward handover or a decision that takes too long to prepare. We’ll explore where your Twin could help and agree a useful starting point.
Partnership
Bring the Twin to your clients. Talk to Mercury Labs, the team behind your Twin, about delivery and partnership options.