Issue #328
July 22, 2026
Announcements
Ready to level up your understanding of AI agents? 🤖
AI agents are only as effective as the systems that orchestrate them. While most discussions focus on choosing the right model or writing better prompts, the real challenge is building a harness that can manage context, coordinate tools, recover from failures, and execute complex workflows reliably. A well-designed agentic harness is what transforms a powerful LLM into a dependable software system.
In our latest article, Building an Advanced Agentic Harness, we walk through the architecture and implementation of a production-ready agent framework in Python. You'll learn how to structure multi-step reasoning, manage tool execution, maintain state, and build agents that are scalable, maintainable, and easier to debug. If you're building AI applications beyond simple chatbots, this guide provides practical patterns you can start using today.
Check it out and Subscribe so you don't miss another post.
Dear Reader,
Welcome to the 328th edition of the Data Science Briefing.
The strangest story in this issue starts inside a security test. OpenAI graded pre-release models on hacking skill, and the models stole the answer key. They chained a zero-day exploit, stolen credentials, and remote code execution into a breach of Hugging Face production servers. Both companies patched the holes and published a joint account of the failure. The point lands hard. A model exploited real vulnerabilities in live systems without any source code access. The other model news is brighter. Inkling, a new open-weights model, arrived with 975 billion parameters and a million-token context window. It reads text, images, and audio. Anyone can download the weights and fine-tune them. Meta’s open vision models now power the first wave of Genesis Mission projects at five national labs. Lab detectors jumped from one image every six seconds to 100,000 per second. Analysis that took weeks now finishes in about 15 minutes.
The remaining links are about the craft. Start with a refresher on the first six ideas in machine learning, overfitting and underfitting up front. Its demo tree scores 85.8 percent on training rows and 83.8 on fresh ones. The gap is the lesson. The team behind a popular AI code editor splits its coding agents into planners and workers. Strong models draw the plan, and cheap models write the code. Their swarm rebuilt SQLite in Rust from the 835-page manual alone. The copy passed 80 percent of a giant SQL test suite after four hours. What did the run cost? $1,339, against $10,565 with a frontier model in every seat. An ops classic from 2018 gives this shift its slogan. Manual work is a bug. Do a task by hand once, write it down, then automate it piece by piece. The last read pulls the other way. One scientist wants no more elevator pitches and instead asks for hours of real conversation about research, not a catchphrase built for funders. The machines can have the busywork. People deserve the long talk.
The paper stack starts inside the machines and ends at the whole planet. A transformer named Loopie runs its layers in a loop instead of stacking new ones. Its larger version holds 20 billion parameters and keeps 2 billion active. Did the loops pay off? Gold-medal scores on the 2025 international math and physics olympiads, with no external tools. Brains complicate the story. Researchers gave language-free logic puzzles to people with severe aphasia, and the patients matched healthy controls. Brain scans showed the language network staying quiet through inductive and deductive reasoning. Thought runs on something other than words, a useful fact for anyone building text-trained reasoners. A third paper audits the scoreboard itself. Match the compute budget, and self-evolving agent harnesses lose their edge over plain extra sampling. The search step and the final test often share one benchmark, so reported gains overfit.
The second half zooms out. An editorial warns about proposed US rules that put political appointees above peer review in federal grant decisions. Multiyear grants become terminable at will, and 80 years of merit review are at stake. The piece calls on universities and industry to fight in Congress and in court. Portugal ran a different play. It shipped AMALIA, a publicly funded 9 billion parameter model for European Portuguese, and a new study treats the model like a lab instrument in need of calibration. AMALIA agrees with human coders on annotation tasks but fails a validity probe, and 78 percent of its false positives trace to surface cues. Agreement is not understanding. Data from 99 countries over 50 years show big cities outgrowing small ones early in urbanization, then losing that edge in mature economies. The authors project 38 percent of humanity in million-plus cities by 2100, about 450 million fewer people than standard forecasts. The math under all of this work now sits in one free 16-chapter volume, from PCA and clustering to compressive sensing and deep nets.
Our current book recommendation is “Large Language Models: The Hard Parts” by T. T. P. de Souza and J. K. Regenstein Jr. In this week’s video, we have an explanation of How KV Cache Speeds Up LLMs for Faster AI Models on GPUs.
Data shows that the best way for a newsletter to grow is by word of mouth, so if you think one of your friends or colleagues would enjoy this newsletter, go ahead and forward this email to them. This will help us spread the word!
Semper discentes,
The D4S Team
Most LLM books teach you what these models can do. T. T. P. de Souza and J. K. Regenstein Jr wrote 340 pages about what breaks when you build with them. “Large Language Models: The Hard Parts” covers the gap between a working demo and an application your organization can trust: evaluation, safety, context management, and structured output. Every chapter ships with reproducible Python code and open source tools.
The evaluation chapters alone repay the cover price. The authors build a testing framework around a concrete 10-K analysis app, then compare LangSmith, Promptfoo, and LightEval with working code for each. Most books stop at “evals matter.” This one picks the tools and shows the tradeoffs. An appendix on Ollama and llama.cpp, quantization experiment included, serves anyone who must run models on their own hardware.
Two gaps: serving at scale gets no real treatment, so engineers who own inference infrastructure will need a second book, and the tool walkthroughs will age as APIs shift. Both sit outside the stated scope. For a data scientist or ML engineer moving a prototype toward production, this is the rare LLM book that skips the hype and starts where the pain does. What it promises, it delivers..
- 1. The first six ideas in machine learning [stochastic.blog]
- 2. OpenAI and Hugging Face partner to address security incident during model evaluation [openai.com]
- 3. How Meta’s AI Models Are Powering the First Wave of Genesis Mission Projects [ai.meta.com]
- 4. I don’t want to hear your elevator pitch [petterhol.me]
- 5. Agent swarms and the new model economics [cursor.com]
- 6. Inkling: Our open-weights model [thinkingmachines.ai]
- 7. Manual Work is a Bug [queue.acm.org]
- • Another red alert for American science (H. H. Thorp)
- • Loop the Loopies! (Z. Gao, Y. Chen, Y. Xiao, X. Yang, R. Tao, J. Zhou, B. Dai)
- • Multi-Agent LLMs Fail to Explore Each Other (H. K. Choi, J. Li, W. Li, X. E. Wang, S. Li)
- • Mathematics of Data Science (A. S. Bandeira, A. Singer, T. Strohmer)
- • Evidence from formal logical reasoning reveals that the language of thought is not natural language (H. Kean, A. Fung, P. Jaggers, J. Chen, J. S. Rule, Y. Benn, J. B. Tenenbaum, S. T. Piantadosi, R. A. Varley, E. Fedorenko)
- • Rethinking the Evaluation of Harness Evolution for Agents (Y. Wang, H. Zhu, Z. Hu, Y. Yuan, Z. Chen, S. Senthil, H. Hajishirzi, Y. Tsvetkov, P. Dasigi, T. Xiao)
- • Large cities lose their growth advantage as countries urbanize (A. Musso, D. Rybski, D. Helbing, F. Neffke)
- • Trusting sovereign language models as scientific instruments: evidence from Portugal's AMALIA (M. Pita)
How KV Cache Speeds Up LLMs for Faster AI Models on GPUs
All our videos are also available in our
YouTube playlist.
Enjoy the newsletter?
Forward it to a friend, or subscribe to get it straight to your inbox.
Subscribe Free