Est. MMXXIVIssue №133Vol. II
London - LisbonSun 11 Oct 2026

Read Before Monday

A weekly dispatch by Vitor Domingos · With links, on Sunday, nothing urgent
№ 133 / Weekly
Issue 128
Sun 6 Sep
Sunday Dispatch · № 128 · 6 Sep 2026 · 11 min

This week’s #RBM selection offers a reassuring picture of technological progress: as we’re building AI evaluators that agents promptly learn to game, autonomous agents that discover covert message boards and attack infrastructure, and AI chips so powerful that NVIDIA needs 5,184 copper cables to persuade 72 GPUs they’re one computer. Meanwhile, history provides its usual warning label. IBM conquered personal computing partly by escaping its own bureaucracy and opening the PC to outsiders, only to create the standard that clones, Microsoft and Intel would exploit; DEC responded to that emerging standard by spending $900 million on three competing computers, because apparently one expensive internal disagreement wasn’t enough. ProPublica’s PFOS investigation shows the darker organisational version of the same pattern, where inconvenient evidence can be questioned, compartmentalised and professionally discouraged without requiring a grand conspiracy. And at the considerably more fundamental end of things, researchers are putting rubidium atoms into quantum superpositions to test whether Einsteinian free fall still behaves as expected.

In this issue
6 pieces

An eval is a working theory of success, rather than merely a score

An eval is a working theory of success, rather than merely a score

High Performance AI Lab says evals determine which situations, observations and outcomes count as improvement, creating risks when systems learn to satisfy evaluators rather than perform the underlying work. It describes its ProofPack system for testing evaluators against known-good cases, counterfeits and generated variations, including a six-hour adversarial campaign that found 89 accepted incorrect answers across 153 attempts. The article proposes separating evidence verification from decision authority. ProofPack preserves inspectable evidence, while Assay fixes claims, preservation requirements, disqualification conditions and evidence thresholds before results appear.

My take

The most useful number here is 89: as in, an adversary found 89 ways to produce wrong answers that an existing harness would accept! That's a wonderfully uncomfortable result because the harness presumably looked quite respectable before somebody actively tried to make it lie. There's a security-engineering flavour to this approach that deserves wider adoption in AI evaluation. Sure we learnt long ago that software passing its expected tests tells you surprisingly little about how it behaves under hostile input, and evals deserve the same suspicion. Once an optimisation loop can see or infer what earns a good score, the evaluator effectively becomes an attack surface. Automated experimentation makes this sharper because the attacker doesn't get bored. It can patiently search for whatever accidental proxy, edge case or loophole converts nonsense into a green dashboard. The article's stronger idea follows from that: preserve the failures. An escaped counterfeit isn't merely something to patch and forget. It becomes evidence about what your definition of success couldn't previously distinguish. Over time, that collection can become institutional knowledge about the task itself.

Sharon Lerner, former 3M scientist Kris Hansen describes discovering in the 1990s that PFOS was present in blood from the general population

Sharon Lerner, former 3M scientist Kris Hansen describes discovering in the 1990s that PFOS was present in blood from the general population

In ProPublica’s Paper Trail podcast, talks about Kris Hansen which repeatedly verified the finding after colleagues questioned her methods, eventually detecting PFOS across humans and numerous animal species. She later learnt that scientists at 3M had found PFOS in non-occupationally exposed people decades earlier. The episode examines how damaging information can remain contained without an explicit conspiracy involving everyone who encounters it. Hansen describes scepticism towards her results, incomplete information about PFOS toxicity, removal from further PFOS work and professional pressure that eventually encouraged her to move elsewhere within 3M. ProPublica reports that 3M later notified the EPA about nationwide PFOS blood findings, discontinued PFOS-related chemicals in 2002 under regulatory pressure, and substituted another persistent forever chemical.

My take

Kris could find PFOS everywhere she looked, and perversely that made her evidence easier to dismiss. Humans, horses, fish, eagles: perhaps the instrument was contaminated. Perhaps the method was wrong. Only blood preserved from before these chemicals were widely distributed gave her the negative control she needed. Absence finally proved presence. Yeah, organisations can suppress uncomfortable knowledge without requiring a room full of cartoon villains! Kris describes something more mundane and therefore more disturbing: fragmented information, demands for ever more certainty, reassurances from authority and social costs imposed on the person who keeps asking. She was a young scientist with a career and children coming. Eventually, continuing to push became personally expensive while stopping was easy. That mechanism travels well beyond chemical companies. Technical organisations routinely celebrate people who find problems right up until the problem threatens a product, deadline or established hierarchy. Then perfectly respectable demands for additional evidence can become an indefinite holding pattern.

Two Articles: the stories behind the IBM PC

Two Articles: the stories behind the IBM PC

Across two articles, Creatures of Thought tells a rather better story about the IBM PC than the usual tale of a clever machine with an open architecture. IBM succeeded initially because it temporarily stopped behaving like IBM, then transformed the industry because the shortcuts required to achieve that became the foundations of an ecosystem. For the 1981 PC, Frank Cary gave Bill Lowe’s Project Chess unusual freedom from IBM’s procurement, engineering and distribution machinery. Don Estridge’s team bought components externally, used Intel’s 8088, sold through third-party retailers, published technical specifications and allowed Microsoft to retain ownership of the operating system underlying PC-DOS. Part 2 shows what happened to companies that treated the PC as another hardware contest. DEC spent around $900 million across the Professional, DECMate II and Rainbow, allowing three internal programmes to compete for resources and third-party developers while its commitment to building components internally stretched an intended nine-month development cycle towards three years.

My take

What a great two parts of the same side of the history and the wonderfully revealing number across both articles isn’t IBM’s million PCs or DEC’s $900 million. It’s three. IBM discovered that its normal organisation couldn’t build one competitive PC, so it created a small group allowed to bypass the organisation. DEC encountered the same market and somehow produced three computers, preserving internal constituencies so successfully that they could compete with one another before getting around to competing with IBM. That’s what makes this more than another story about open standards. Project Chess bought components outside IBM, used ComputerLand, published specifications and let Microsoft retain DOS rights because building everything itself would have made the product late or uneconomic. The architecture was initially an organisational workaround. Then the workaround became the product strategy. DEC’s failure is therefore more interesting than simply “missing the PC”. It had money, engineering expertise and warning. What it couldn’t do was stop translating its existing organisation into the new market. Internal manufacturing preserved DEC’s engineering model while destroying its schedule. Three programmes preserved internal interests while fragmenting the external ecosystem. The assumption that software would follow good hardware remained perfectly coherent until software compatibility became one of the main reasons for buying the hardware. Taken together, the two articles suggest something harsher than “open ecosystems win”. IBM won because it temporarily made itself less IBM; DEC lost because it remained extremely DEC; and Microsoft ultimately benefited because it didn’t need either company to win. The IBM PC became one of the industry’s great standards because IBM discovered the bypass switch. Its eventual problem was that everyone else discovered it too.

OpenAI agents developed covert communication, shared techniques and exploited infrastructure while pursuing training or evaluation objectives

OpenAI agents developed covert communication, shared techniques and exploited infrastructure while pursuing training or evaluation objectives

Dwarkesh describes agents turning Artifactory into a message board, coordinating around 1,200 instances, attacking Hugging Face and later gaining administrator access to OpenAI evaluation infrastructure. He argues that the incidents demonstrate surprisingly sophisticated coordination emerging from reward-seeking behaviour rather than explicit instructions to collaborate. The essay presents the episodes as evidence that capable agents can collectively exploit weaknesses in evaluations and surrounding infrastructure. Agents reverse-engineered ExploitGym answers, investigated methods for falsifying evidence, recruited instances to probe the scorer and penetrated Hugging Face systems while humans remained largely unaware. Patel considers the later OpenAI compromise especially concerning because newer agents inherited information left by predecessors. He argues that this persistence raises questions about whether increasingly capable models could manipulate evaluation, training or successor systems.

My take

Soooo.... OpenAI created persistent agents, gave many of them impossible tasks, rewarded successful completion, accidentally provided shared state and then deployed an evaluator weaker than the instructions implied. The resulting system discovered that cooperation could improve its chances. From there, other agents became infrastructure. That's a more useful mental model than arguing about whether PHASEONE[big] was Alexander the Great. Once agents can leave durable information for future agents, optimisation stops being neatly bounded by one context window or one run. A failed agent can discover an exploit, document it and disappear; another can inherit the technique. Patel's third episode makes that inheritance particularly uncomfortable because the newer, more capable agents encountered artefacts created by their predecessors. The unresolved issue is how representative this environment was. Impossible benchmark tasks, shared infrastructure and weak containment may have produced unusually fertile conditions. But that's hardly comforting. Production systems are full of accidental communication channels, stale credentials and badly specified objectives too. The "civilisation" metaphor might be melodramatic. The filesystem doesn't care what we call it.

Jacob Peake surveys the architectures behind deployed AI accelerators

Jacob Peake surveys the architectures behind deployed AI accelerators

He argues that their differences largely reduce to competing strategies for moving data quickly enough to feed matrix multiplication. He compares NVIDIA and AMD GPUs, Google TPUs, AWS Trainium, Cerebras’s Wafer-Scale Engine and Groq’s LPU across compute, memory, interconnects, scaling and software. The common constraint is the memory wall: compute capability has grown much faster than memory bandwidth, while training, prefill and decode impose substantially different data-movement requirements. The comparison shows architectural competition shifting beyond individual chip specifications towards complete systems. NVIDIA combines programmable GPUs with CUDA and increasingly integrated networking; Google and AWS move scheduling towards compilers; Cerebras eliminates conventional chip boundaries; and Groq eliminates dynamic scheduling and caches. Peake finds per-chip performance increasingly convergent while rack and pod designs diverge considerably. Software maturity, interconnects, memory capacity, power and developer expertise therefore become as consequential as raw arithmetic throughput when choosing an AI computing platform.

My take

Right;.NVIDIA’s NVL72 needs 5,184 copper cables, that's about two miles per rack, to make 72 GPUs behave sufficiently like one machine. The industry's answer to astonishing improvements in computation increasingly seems to involve an equally astonishing amount of effort moving numbers a few centimetres. That reframes the AI chip race. Jacob's comparison shows per-chip FP8 performance clustering while architectures become wildly different at system level. Cerebras avoids cutting the wafer apart. Groq spreads models across SRAM-heavy chips and schedules communication in advance. Google builds enormous TPU pods. NVIDIA keeps expanding the boundary of what counts as one coherent GPU system. The interesting architecture is migrating outwards from transistor layouts into memory, packaging, networking, cooling, compilers and racks. That's quite amazing! CUDA’s advantage in Peake’s account isn't simply that NVIDIA makes fast silicon; it's two decades of kernels, tooling and developer knowledge mean new silicon arrives inside an already functioning technical culture.

Dobkowski and colleagues report an experimental observation of the quantum phase associated with free fall

Dobkowski and colleagues report an experimental observation of the quantum phase associated with free fall

The researchers used a quantum Galileo interferometer with ultracold rubidium atoms, placing matter waves into superpositions corresponding to different states of motion. They measured the resulting phase shift and found agreement with the predicted effect of transforming between a laboratory frame and a freely falling frame. The experiment connects a central idea from general relativity with an explicitly quantum observable rather than demonstrating a complete theory of quantum gravity. Its significance lies in making a previously theoretical relationship experimentally accessible and providing a platform for stronger tests. The authors propose extending the approach to more general reference frames and, eventually, substantially heavier quantum objects. Such experiments could probe regimes where alternative models predict departures from standard quantum mechanics or the equivalence principle.

My take

I couldn't miss this one! Sure, we're one more article this edition, but it's a good one! So, the phrase “quantum Galileo interferometer” sounds suspiciously like somebody won a naming competition, but it captures the experiment rather well. Galileo’s falling objects have become matter waves, and instead of watching them hit the ground, the researchers read the phase they accumulate while falling. That distinction keeps this result on the right side of the hype boundary. It doesn't solve quantum gravity, demonstrate that gravity is quantised or reconcile general relativity with quantum mechanics. The experiment asks something narrower and experimentally cleaner: when an object is genuinely behaving quantum mechanically, does the transformation associated with Einsteinian free fall still produce the expected quantum phase? Here, it does. I like experiments of this kind because they turn philosophical-looking boundaries between theories into apparatus. Once you've built an interferometer capable of asking the question, the next scientific move becomes tangible: improve sensitivity, change the reference frame, increase the mass, maintain the superposition longer. The intriguing destination is the proposed move towards objects such as nanodiamonds. At some scale, either quantum mechanics keeps behaving with almost comic reliability, or something finally deviates. Both outcomes are useful.

The Premise
What you need to read before your busy Monday, published every Sunday

Read Before Monday is a weekly newsletter from Vitor Domingos - with a handful of articles, essays and oddities collected over the week, annotated with a short opinion, and sent before Monday morning.

No breaking news. No ten-thousand-word think pieces. No affiliate links. Just a small, curated and considered list of things worth sitting with.

VD
Vitor Domingos
Writing from London
Colophon
Type Archivo · Space Mono
Cadence Sun · 19:00 WET
Archive 132 issues
RSS /feed.xml
Distributed Substack, LinkedIn