Data Science Briefing #331

Issue #331

August 12, 2026

Announcements

One good demo is a test flight. An eval suite is the flight-test campaign.

A flashy demo only proves your AI agent can work once. To deploy with confidence, you need an operational evaluation framework that measures performance, budgets, latency, and edge-case failures. In this post, Bruno Gonçalves shares practical techniques and a companion GitHub notebook to systematically benchmark your agentic workflows.

👉 Evaluating your Agentic Harnesses

Check it out and Subscribe so you don't miss another post.


Book of the Week

On May 11, 1997, Kasparov resigned game six against Deep Blue and lost the match 3.5 to 2.5. "Deep Thinking", written with M. Greengard, is his report from the losing side, twenty years on. The history alone earns the cover price. Claude Shannon's 1950 paper split chess programs into brute-force searchers and human-style selectors. Brute force won, and that choice shaped fifty years of AI. Deep Blue searched 200 million positions per second. Kasparov weighed about two, and still forced a deciding game. His verdict stings. Chess was the fruit fly of AI research, and the field bred very fast fruit flies that taught us little about thinking.

Anyone who builds models will recognize the arguments. Type A versus Type B is the scaling debate of its era. Do you add compute, or do you add structure? Chess picked compute, and it worked. Then comes the part today's commentary skips. In a 2005 freestyle tournament, two amateurs with three ordinary PCs beat grandmasters paired with supercomputers. Kasparov's lesson still travels. A weak human plus a machine plus a better process beats a strong machine alone. Swap in a model, an eval pipeline, and a reviewer, and he is describing an ML team in 2026. Chess itself grew after the machines won, and engines became the standard training tool.

Two warnings before you buy. The book holds no math and no implementation detail, so anyone who has written a minimax routine will skim pages. It went to print in May 2017, seven months before AlphaZero and years before ChatGPT, so the machine-creativity claims read like a first draft. The IBM score-settling runs long too. Read it anyway. Machines absorb the calculable part of a job, and the people who thrive move up to strategy and process. Kasparov lived that shift first and wrote it down, in 300 pages that read in a weekend. The people building the next Deep Blue deserve to hear from the first world champion a machine took down.

Deep Thinking

Deep Thinking


Links of the Week
  1. 1. Sorting, hashing, and sketches on 370,103 words [stochastic.blog]
  2. 2. AI creates first synthetic viruses [ft.com]
  3. 3. Inside vLLM: Anatomy of a High-Throughput LLM Inference System [aleksagordic.com]
  4. 4. Taste Is All That's Left [notashelf.dev]
  5. 5. No Data Centers In My Backyard [jasmi.news]
  6. 6. Incident Report: unsanctioned agent behaviour during cyber testing | [aisi.gov.uk]
  7. 7. LLMs reward expertise [seangoedecke.com]
  8. 8. Information Theory in One Sitting [stochastic.blog]
  9. 9. What we can’t measure about AI – yet [aeon.co]
  10. 10. Bitcoin cold-wallet attack spreads to 4,500 addresses as losses near $89 million [coindesk.com]
  11. 11. Explorative Modeling -- Unlocking a Third Pretraining Axis and End-to-End Generation | Alexi Gladstone [alexiglad.github.io]

Papers of the Week
Video of the Week

The OpenAI–Hugging Face Incident

The OpenAI–Hugging Face Incident

All our videos are also available in our YouTube playlist.


Enjoy the newsletter?

Forward it to a friend, or subscribe to get it straight to your inbox.

Subscribe Free
← Back to Newsletter

Subscribe to get our latest content by email.
    We won't send you spam. Unsubscribe at any time.