
This week’s #RBM reading was mostly about technology discovering, once again, that the difficult bit is rarely the bit on the brochure. We start with Denver Airport that built nearly 20 miles of automated baggage track, 4,000 Telecars and a small army of 486s, only to discover that moving bags was comparatively easy; having an empty cart in precisely the right place at precisely the right time was the actual computer science problem. Then, a history of computer memory makes much the same point from another direction: programmability gets the glamour, while generations of engineers quietly made bits race through mercury. Drew Breunig’s Audio AR takes us into the present, where headphones, speech models and contextual assistants can finally whisper useful information into our ears, provided they first solve the rather important problem of knowing when to shut up. Education has its own version of this recurring enthusiasm: instructional television, MOOCs, smart boards, open gradebooks and now AI have all arrived carrying some promise of efficiency. Richard Cook’s How Complex Systems Fail, written more than a quarter of a century ago, supplies the manual for all of this: complex systems are already full of faults, humans constantly compensate for them, disasters emerge when several normally survivable problems align, and afterwards everyone goes looking for the reassuringly simple “root cause”. See you next week!
In this issue
5 piecesJ. B. Crawford traces Denver International Airport’s automated baggage system from ambitious design to operational failure

The article argues that the decisive problem was not a single bad algorithm or machine, but a tightly coupled physical and software system rushed into airport-wide use. Evidence includes nearly 20 miles of track, almost 4,000 Telecars, over 100 486 computers, reactive empty-cart dispatch, unreliable barcode associations and mechanical failures that cascaded into queues. Crawford also complicates the usual morality tale. BAE reportedly warned Denver that the schedule was too short, requirements kept changing, and the small demonstration system never exercised the hard dispatch problem. Denver eventually installed a roughly $50 million conventional backup and opened 16 months late, while United kept a reduced Telecar system until 2005. The later success of similar individual-carrier technology in Orlando suggests the concept was viable, but Denver’s scale, timetable and integration discipline were not.
So, the phrase I keep coming back to is “empty cart management system”. It is wonderfully unglamorous and it explains why the project is more useful than the usual story about flying suitcases... Sure, the visible job was moving bags, but the real hard job was continually putting spare capacity in the right place before demand arrived. Once that failed, every loading point became a queue whose failure could spread elsewhere. That pattern turns up whenever automation removes local slack. A human with a tug can improvise, wait, reroute or prioritise. Denver Airport replaced much of that discretion with a network whose physical state, database state and schedule assumptions all had to agree in real time. Lose a barcode association, misplace an empty cart or take a corner too fast, and the software’s model of the airport starts drifting away from the airport itself. The system then spends its effort recovering from yesterday’s mistakes while new bags keep arriving. This showed that the machinery could perform its tricks, while neatly avoiding the shared-resource problem that defined production. That is a familiar way to prove the easy part and call it validation. Denver’s later fallback, belts and human-driven tugs squeezed through tunnels with railway-style signalling, looks primitive only until you notice that it actually opened the airport.
Electronically controlled working memory, rather than programmable machinery alone, was the breakthrough that made modern computers practical

lcamtuf’s article argues that Mechanical calculators and programmable devices already existed, but moving intermediate results around complex machines was unwieldy. It traces working memory from relay and vacuum-tube latches through capacitor drums, acoustic delay lines, Williams and Selectron tubes, magnetic drums and core memory, showing how each technology traded speed, density, cost and complexity. The article separates working memory from bulk storage, a distinction it argues is often blurred in popular computing history. Early computers frequently avoided expensive electronic latches by storing information in transient physical phenomena that required continual refreshing or circulation. Magnetic core memory later offered faster, compact random access before integrated circuits displaced it. The article connects these designs to modern SRAM and DRAM, while arguing that RAM capacity remained a major constraint on personal computing until the late twentieth century.
The image that sticks is that memory as something literally moving! Its bits racing through mercury as sound, twisting around coils of wire, circling on magnetic drums. We now talk about memory as though it were a place, an addressable landscape of neat little cells. Much of its history looks more like choreography. So that changes how I read the article’s argument about Babbage; programmability is intellectually glamorous because it sounds like software, yet working memory sounds like plumbing. But once computation becomes complicated, retaining an intermediate result and getting it back to the right place quickly is what lets complexity accumulate rather than collapse under its own mechanics. The bizarre hardware here was solving that coordination problem with whatever physics happened to be affordable. The continuity with DRAM is especially good. A modern capacitor cell is vastly removed from a rotating drum of mechanically switched capacitors, yet both accept impermanence as an engineering bargain: make storage cheap and dense, then spend effort continually rescuing the data before physics erases it. Bottom line; computer architecture has repeatedly advanced by finding tolerable ways to make unreliable or awkward physical phenomena behave like dependable abstractions, but memory only looks passive after an extraordinary amount of machinery has made us forget what it is doing.
Audio Augmented Reality is becoming practical because smart headphones

Drew Breunig argues that better speech recognition and synthesis, large language models, and richer contextual data can deliver information without a screen. Earlier systems such as Microsoft Soundscape and Foursquare’s MarsBot exposed the central UX problem: useful audio depends on surfacing the right information at the right moment, while too many irrelevant alerts quickly become intrusive. Drew contrasts three interaction patterns: VoiceMap’s route-bound audio tours, Meta’s pull-based Wayfarer glasses, and Apple Fitness’s opt-in contextual sessions. He argues that constrained sessions work because applications know what the user is doing and can push relevant information with greater confidence. More general Audio AR would require better access to calendars, messages, location and other contextual signals, plus open standards for headphones, context sharing and voice assistants.
Luther Stickell is useful because he speaks sparingly; he has context, but more importantly he has judgement about interruption. That turns Audio AR from a speech-interface problem into an attention-allocation problem. Drew’s three product patterns expose a useful gradient. VoiceMap works because the route already encodes intent. Apple Fitness works because starting a workout creates a bounded session. Meta’s pull model works because the user chooses the moment. Each product gets better as it reduces uncertainty about whether an interruption is welcome. The more ambient the system becomes, the less certainty it has and the more damage a badly timed cue can do. That makes the call for richer context necessary and slightly dangerous from a product-design perspective. More signals may improve relevance, but they don’t automatically produce judgement. A model can know my location, calendar and messages and still be annoyingly literal. The hard engineering target is probably not maximal context. It is confidence calibration: know enough to help, and be conservative enough to stay silent when uncertain. An assistant that occasionally misses a useful cue is tolerable. One that keeps guessing wrong becomes something you switch off.
Ivan Johnson argues that education repeatedly mistakes technology for a way to make teaching cheaper

He uses instructional television, MOOCs, smart boards and open gradebooks as successive examples, including his own decision to spend more than $50,000 installing smart boards that teachers resisted and largely abandoned. Johnson’s central distinction is between technology that extends teacher capacity and technology that substitutes for teacher judgement. He says tools such as photocopiers, Google Read&Write and AI-assisted lesson differentiation can return time to teachers, while systems that decide what or how to teach risk turning adults into monitors. His open-gradebook example supports the broader warning: constant visibility allegedly shifted attention towards points, increased parental pressure and discouraged low-stakes work. The article ultimately defends education’s labour intensity as a feature rather than an inefficiency.
The part I’d keep is the open gradebook, because it shows how a “transparent” interface can quietly redesign the system around it. A grade used to arrive at intervals. Make it continuously visible and it becomes a live metric, which invites continuous checking, parental escalation and optimisation by everyone being measured. Johnson’s 50-to-80-times-a-day figure is methodologically thin, but the mechanism deserves attention. That is more interesting than another argument about whether AI will replace teachers. The useful distinction is who retains discretion when software enters the loop. A photocopier speeds up production without deciding what deserves copying. An open gradebook exposes a metric and changes behaviour. AI can go further by generating material, sequencing work and recommending interventions, so the design question is not simply whether a teacher remains physically present. Once a proxy becomes constantly observable, people start managing the proxy. Schools are hardly unique there. Anyone who has watched a dashboard KPI become a target should recognise the smell.
18 propositions about failure in hazardous, highly defended systems such as healthcare, transport and power generation

Richard I. Cook argues that catastrophes rarely come from a single broken component or bad decision. They emerge when multiple small, normally contained failures align, while operators work inside systems that already contain latent flaws and degraded components. The treatise therefore rejects isolated “root cause” explanations as a poor model of systemic accidents. Cook also argues that people are central to safety because practitioners continually adapt operations, balance production against risk, and block failure trajectories that formal defences miss. Hindsight then makes those same decisions look obviously mistaken after an accident. Changes intended to remove familiar problems can introduce rarer, higher-consequence failure modes, while narrowly targeted post-accident fixes may add coupling and complexity. Safety, in this account, is an emergent and continuously recreated property of the whole system.
Cook’s claim that complex systems “run as broken systems” is quite true! So it is a useful corrective to the tidy picture produced by dashboards and postmortems. A service can be green because the machinery is healthy, or because people are quietly rerouting traffic, restarting jobs, carrying tribal knowledge and compensating for defects. The observable outcome is the same until those compensations stop lining up. For engineering leaders, that makes uneventful operation ambiguous evidence. If reliability depends on expert improvisation, absence of incidents may conceal fragility rather than demonstrate it. Cook’s “proto-accidents” are therefore more interesting than the final outage: the near miss, manual recovery or weird alert that disappeared after someone nudged a system back into shape. Those moments expose the operating envelope while everyone still remembers what they did. Waiting for catastrophe selects only the rare combinations that escaped every defence, then tempts us to narrate them backwards into inevitability. A system’s resilience may be most visible on the days when nothing officially happened.
My Week in AI
- OpenAI bots meddled with US government agencies, including SEC and Census | bbc.co.uk
- Revealing the details of how OpenAI agents hacked Hugging Face | swarmtraces.org
- Plan mode is dead | aymannadeem.com
- AI labs need to start funding historical research | resobscura.substack.com
- Can gzip be a language model? | nathan.rs
- LLM Transformer Model Visually Explained | poloclub.github.io