FR

About

Theo Martin, portrait

9 years of putting MLMachine learning: programs that learn a rule from examples instead of getting it written by hand. and then LLMsLarge language model: a model trained on huge amounts of text to predict what comes next in a text. Claude, GPT and Gemini are some. into production, first on pricing at Amazon, then on product catalogs in SaaS. Today I use coding agents every day, at a rate of >100B tokensThe pieces of text, a few letters each, that a model reads and writes. It is the unit you pay for., and I work with teams that want to get more out of them: developers, product managers, designers.

The path

The same path, in pixels

I started at Amazon, in Luxembourg: 3.5 years on European pricing. I came in through supply chainPurchasing, stock, warehouses and deliveries. optimization and then AWS architecture, with 4 AWS certifications passed in 4 months, including the Solutions Architect ProfessionalThe highest AWS certification for designing cloud architectures.. in pixels

Then pricing, as the only data scientist on a team that went from 3 to 13 people. I did causal inferenceMeasuring the real effect of a decision (a price drop, for example) by separating it from everything else that changed at the same time., with synthetic controlA method that builds a “twin” of what you changed from a mix of things you did not change, then compares the two. coded by hand before the libraries were mature, hierarchical Bayesian modelsStatistical models that share information between close groups: a product with few sales borrows what is known about its category. of price elasticityHow much sales move when the price moves by 1%., and hedonic pricesExplaining a price by the value of each feature of the product: size, brand, memory, and so on. estimated on text and image embeddingsA list of numbers that represents a text or an image, built so that two close things get close numbers., with ELMoA 2018 language model, one of the first to represent a word according to the sentence it is in. and then BERTThe language model Google published in late 2018, the base of most text-processing systems until LLMs arrived. fine-tunedTrained a bit more on your own data, to specialize a model that is already trained. in 2019. in pixels

Out of that came the price volatility monitoring program, which I started and which ended up adopted across the whole group worldwide. 30B price changes analyzed, an anomaly caught across 1B visits. in pixels

Then Unifai, where I was the company’s first ML engineer. An end to end MLOpsEverything that keeps a model running in production: data collection, training, deployment, monitoring. pipeline on GCPGoogle Cloud Platform, Google’s cloud. for retail groups, and language models from before ChatGPT, FLAN-T5An open language model from Google, released in late 2022. in zero-shotUsing a model on a task without showing it a single example. against my fine-tuned extractor, on real industrial catalogs: better on booleans, worse on numbers. That work is what put the company in a position to be bought: during the due diligenceThe audit a buyer runs on a company before buying it., I walked the investors through the architecture and the ML infrastructure, up to the acquisition by Akeneo in 2023. in pixels

Then 2.5 years at Akeneo as Tech Lead of the Core AI team. I architected and ran the internal inference platformThe service that receives requests from other teams and sends them to the models, handling quotas, errors and costs. that every product team depended on, and behind them >100 enterprise retailers. I wrote the library that sends that traffic to VertexAI, OpenAI and Anthropic through a single LiteLLM proxyAn open source intermediate server that gives one single interface to call all model vendors., with each customer’s cost tracked in the response headersThe metadata a server sends back with each response, next to the content itself., and the log volume per request dropped by 50 to 75%. And the Data Architect Agent, a multi-agent systemSeveral agents that split a task, each with its own role, and hand over to each other. with human gatesSteps where the system stops and waits for a human to approve before going on. that took catalog onboarding for an enterprise retailer from several months down to a few days. in pixels

From the tool to the harness

At Akeneo, I pushed Cursor in my team, then Claude Code as soon as it came out. I gave classes to several teams, and people I had convinced then convinced their own teams. I also pushed the leadership, very early and very hard, to pay for licences: in the end everyone got their own for Claude Code. in pixels

Then I went from the tool to the harnessEverything you put around the model to make it an agent: instructions, tools, checks, permissions.: CLAUDE.mdThe instruction file that Claude Code reads at startup in a project. The open equivalent is called AGENTS.md. files versioned in the repoThe folder that holds a project’s code and all its change history, managed with git., so that everyone’s harness gets better without each person having to look after their own. My setup ended up as the base of an internal training.

What I think

Positions reviewed on October 6, 2026.

The first win stops many teams, and so does the first failure

I have never yet seen a team where the tool was the limiting factor. Often the team installs the agent, sees it succeed at one task and thinks “ok, it can do that”, or sees it fail and stops there. In both cases it believes it has found the limits of the agent, when it was lucky, or not.

When code does not compile, you know it is the code. With an agent, it is hard to tell the limit of our usage from the limit of the tool, all the more because the result depends on the context you give it. When a dev at another company says they wrote the same function, you know it is doable and that the problem is on your side. When they say their agent does it, you think it depends on their setup.

What I aim for is a team that really delegates. On my side, the agent writes whole features in an autonomous loopThe agent chains code, tests and fixes on its own until the result passes, without waiting for a human at each step. from Figma mockups, leaves its questions as comments in Notion, and picks the work back up as soon as a PM has answered.

Models and harnesses are converging

The lead changes hands within days. Since January 2025, Claude has had the best model 55% of the time, GPT 45%. In September 2026, Fable 5.1, GPT-6 Astra and then Opus 5.5 each went ahead on Terminal-BenchA public benchmark that measures whether an agent can finish real tasks in a terminal., within three weeks. A jump forward gets caught up in a few weeks, which is why people talk about frontier modelsThe few most advanced models of the moment, at OpenAI, Anthropic and Google, which stay close to each other.. Harnesses follow the same slope.

Claude and GPT since 2025Oct 2026

Scores published at release: SWE-bench Verified in 2025, from Claude 3.7 Sonnet (63.7%) to Claude Opus 4.5 (80.9%); Terminal-Bench 2.0 in early 2026, up to GPT-5.5 (82.7%); Terminal-Bench 2.1 from May to July, Claude Mythos 5 (88.0%) passes GPT-5.5, then GPT-5.6 Sol (88.8%); Terminal-Bench 4.0 since July 2026, from Claude Opus 5 (52.3%) to Claude Sonnet 5.5 (70.6%), with GPT-6 Astra (57.9%) passing Claude Fable 5.1 (55.8%) two days after it.

What has always mattered is context: what the agent knows about your code and your rules, and what it can check on its own. Once the context is in place, it is worth trying other harnesses and other models. For example have a model’s code reviewed by a competitor: GPT and Opus were a priori not trained on the same data, so the gap between them is much bigger than between Opus and Sonnet, which come out of the same house.

The mental model is worth more than the stack

What works today will be replaced. I built my own automatic review system for PRsPull request: a proposed change to the code, which a teammate reads before it is merged., then I switched to Copilot’s when it came out. A new model no longer needs some of the instructions you had to write for the previous one. What stays is a good idea of what an LLM is, and the ability to adapt.

I read 269 public repos file by file, and it matches what I see in the teams I work with: very few automate the expiry of the files that steer their agents.

What disappears is a trade, not people

In September 2024 I wrote that “80%+ of us” would be obsolete, “probably within 3 years”. I was wrong on one point: it is not the dev who disappears, it is the part of the work that turns a clear spec into code. Taking a ticket and shipping a PR, that part of the cycle is already replaced. It is the work we give to juniors, hence the idea that they go first.

And it moves up the chain. A PM takes customer feedback and the market as input, and produces specs. AI already helps write the spec and sort the feedback, and covers a bit more step by step. Where the chain is continuous, the agent holds it alone. Where it breaks, you put a human.

What AI can do on its ownFeb 2026

The tasks of the product cycle, from completing lines of code to customer feedback, move one by one from humans to AI between January 2021 and February 2026 Product Code Check and ship Align the team Write the spec Split into tickets Read customer feedback Write a feature Write whole files Complete functions Complete lines Write the tests Make the tests pass Review the PR Ship to prod
a human, AI.

The trade moves towards specification, verification and decision. The figures I have seen show a market that closes at the entry level rather than a disappearance.

I do not believe in a “handmade” niche for code. Someone who buys a hand-blown glass vase pays for how it was made, someone who buys software only looks at what it does. Love of the craft for itself saved some artisans, unfortunately it will not save coders.

We are still reviewers, and reviewing guarantees nothing

I work like a director checking on his seniors: result-oriented, with tests the agent has to make pass. Reviewing lowers the probability of error without eradicating it: 3 reviews are better than 1.

With robust tests on isolated parts, it is possible, believe me, not to review some zones. And such a zone can be a whole feature, even a whole piece of software.

40 tickets, 3 ways to ship them

The same 40 tickets, all with tests. A dev coding with AI and no context: 6 bugs in production, 64 human hours. Claude Code with the repo context and one human review: 1 bug, 21 hours. Tests reviewed by a human, Copilot back and forth, human review of critical code only: 0 bugs, 12 hours.
1 hour to code a PR with AI, 30 min per review, 10 min to review the tests.

Here, without the repo context, AI writes twice as many bugs, and the dev spends their days steering it. With the context, the agent codes on its own. In what I set up, the human reviews the tests rather than the code, and only reviews code where it is critical.

The rest is still reviewed: what cannot be tested automatically, tests that are too long or too expensive, and what is critical, like race conditionsBugs that only show up when 2 operations cross in a certain order, so they are rare and hard to reproduce.. AI still gets it a bit wrong, and the whole point is to know where. But I am convinced of it, soon AI will be able to do everything, review included.

Everything that can be tested, I test

My own harness included. I have a test setA fixed series of test tasks, replayed at each change to compare before and after. of 22 tests that I run from time to time, and I check regularly that they still pass. To shrink a CLAUDE.md, I remove everything, see what still passes, then put the rules back one by one. Last time the file went from 108 to 30 lines and all the tests still pass.

Shrinking a CLAUDE.mdStep 1 · The starting file

108 lines at the start. With no rules, some tests fail. Each rule added back makes a few pass again, and at 30 lines they all pass.

What I am working on

I have worked with large companies, one of them in the CAC 40. At a large company, what blocks agent adoption is almost never technical: it is who decides, on what criteria, and when someone writes the decision down somewhere.

I take short engagements, often remote, with teams that want their harness to really pay off. An audit to find out whether your guards hold when the agent switches tools, a workshop day inside your repo with the work committed that same evening, regular office hours to catch what drifts, or a build when what you need is shipped code. The simplest way to start is the audit.

See the services

I also write now and then about what I break and what I fix along the way.

Read the writing


Based in Paris, European time zone, with an overlap with the US East Coast until 5pm ET. Remote anywhere, on site anywhere in France. [email protected], LinkedIn, GitHub.