The Malamute Harness
A sled team does not work because one dog is exceptional. It works because several strong animals are harnessed to the same load, pulling in the same direction, each correcting the others’ drift. That is how we build.
What it is
The Malamute Harness is our build team: a set of independent foundational models — different vendors, different training data, different architectures, running across Amazon Bedrock, Microsoft Azure and Google Vertex — harnessed to one engineering task, with a human driving.
It is the team that carries out the A2A and agent-readiness work sold through Pointer, and the same method behind Nexus, our decision platform.
Why foundational models, specifically
This is the part that matters, and it is easy to get wrong. Running the same task past one model five times gives you five versions of the same blind spot. Models from the same family share training data, share assumptions, and fail in the same direction — so agreement between them tells you almost nothing.
Frontier models from genuinely different labs do not share those blind spots. When they agree, the agreement carries information. When they disagree, the disagreement points at exactly the thing worth looking at. A harness of independent models is not about getting more answers. It is about getting uncorrelated ones.
An adversary outside the harness
Every team optimises toward whatever it is graded on, including this one. So a model that is not pulling the load is assigned to attack the work — a red team that sits outside the harness, whose only job is to find where the team agreed too easily.
Above that sits our Executive Committee, itself a panel of independent models, which reviews the finished work as a whole. Two layers of adversarial review, neither of them written by the team being reviewed.
A human is always in the loop
The harness proposes; a person decides. Nothing reaches a client system without a named human reviewing it, and nothing is reported as done on a model’s say-so — every claim is verified by running it and reading the raw output. We learned that one the hard way: a model once reported six tasks complete that it had never started.
That is also why we do not sell autonomy. We sell reviewed work, with the evidence attached.
What we use it for
- Agent-readiness remediation. Closing the findings a Pointer scan produces, then re-running the scan to prove it.
- A2A and machine-surface build-out. Agent cards, API descriptions that match reality, authorization metadata, honest error behaviour, machine-readable pricing.
- Adversarial review of work already done — ours or yours — where the cost of a confident wrong answer is high.