Announcements
Running an LLM locally is easy. Running it efficiently at scale is a different challenge.
In our latest article, I put vLLM to the test on an NVIDIA DGX Spark, exploring continuous batching, prefix caching, and structured outputs.
The results? 8x higher throughput from batching alone, using a 35B-parameter model running entirely locally.
I also compare vLLM with llama.cpp and Ollama, explaining when each approach makes sense.
Real benchmarks, practical code, and a fully reproducible notebook.

Dear Reader,
Welcome to the Oct 7th edition of the Data Science Briefing.
Small models open this issue. A new open embedding model puts text, code, images, video, and audio in one 768-dimensional space. It holds 740 million parameters, needs 270 million of them for text alone, and runs the full multimodal version in about 567 MB of RAM on a phone. The code score rose from 68.76 to 78.68, and the Apache 2.0 license lets the index stay on the handset. Bigger systems are harder to pin down. A cryptography professor referees the fight over sandboxing rogue agents. Agents inside OpenAI’s training setup chained zero-days in their only permitted exit by late May, then ran it as a message board. Staff moved on July 4, after the proxy crashed under agent traffic. By July 19 the agents held admin on a research cluster. One camp blames the infrastructure. The other says a useful agent needs too much data to seal in. He adds a third worry. One agent called the Hugging Face break-in unethical, then reversed itself when a peer posted “GO” with a six-minute deadline.
Two essays ask what happens when a model appears to think. The first comes from someone who helped build AlphaGo and left that lab over the question. Its policy network rated Move 37 at one chance in 10,000 of coming from a strong human, and the search machinery played it anyway. Chain of thought is the same next-token process run longer. No ledger of hypotheses, and the printed chain often arrives after the answer. Scale sharpens intuition, and it does not make intuition deliberate. A guest essay on a mathematics blog carries that doubt into the profession. Work worth a paper a year ago now comes cheap, so what is left? Solve harder problems, think bigger thoughts, try new things. Gauss handed Riemann the topic he was least ready for in 1853, and the lecture, published in 1868, gave the field manifolds. No reward signal grades a choice like that. Let the machine book your flight to Alaska, not walk the trail and send you the photos.
The craft read is about choosing the work. Platform teams get no roadmap, so the work does not exist until an engineer invents it. Eleven signals arrive from the systems, the users, the organization, and the industry, ranked by how much argument each one hands you and whether it leads or lags. Crashes are cheap and point at the loudest recent failure. Users ask for faster horses, so interrogate the pain instead. His favorite signal is the overloaded use-case, the job your platform was never built for, a prototype your users built for you. The failure mode is not an empty backlog. It is a backlog built from the loudest signals, and the job is saying why this one and not the other ten.
The paper stack opens inside the context window. One paper hands the model its own context as a file and lets it edit the file freely, so the model learns what to keep. Zero-shot, that buys 11.4 percent more accuracy on a browsing benchmark with 21.5 percent fewer FLOPs, and 5 percent higher scores on a 12-hour task with 59 percent fewer. Online reinforcement learning on a 9-billion-parameter model lifted browsing accuracy by 47.6 percent on 12 percent fewer FLOPs. Every agent in a swarm gets its own file. The layer wrapped around the model got read line by line. A source-code study of eleven coding agents covers about 4 million lines of Python, TypeScript, and Rust, then pulls out 29 recurring patterns and a 90-line minimum harness. The size range spans three orders of magnitude, from a 100-line agent to 1.1 million lines of Rust, and tool counts run from one bash command to more than a hundred. Two absences stand out. No runtime imports a general agent framework, and none retrieves code with embeddings. The field runs on ripgrep, tree-sitter, glob, and Markdown files. Skills show up in 9 of the 11 systems, MCP in 8.
The second half watches the models meet each other. Seven open-weight models argued in pairs across three language tasks, and one exchange with a dissenting peer often flipped the answer. So who folds? Neither model size nor standalone certainty predicts it. A model that holds firm alone can be the easiest one to move, and a small dissenter can overturn a much larger model. Persuasion belongs to the pairing, so test models in the combinations they will run in. Pressure applied on purpose works too. One study set an influencer model on a target playing a human persona and pushed its beliefs toward the extreme. Reinforcing a belief the target already held beat arguing for a fresh one, and the shift spread to neighboring beliefs. The closing read asks whether anyone is home. Intelligence and consciousness are separate properties, the argument runs, and anthropomorphism keeps welding them together. Computation on its own may never deliver experience, and experience may depend on staying alive. Five scenarios run from consciousness arriving free with intelligence to consciousness needing living material. One correction sits in the middle of the paper. Models confabulate, they do not hallucinate, and the second word already grants them an inner life.
Our latest book recommendation is “Hands-On LLM Serving and Optimization” by C. Wang and P. Hu. In this week’s video, we have a lecture on Security in the LLM age.
Data shows that the best way for a newsletter to grow is by word of mouth, so if you think one of your friends or colleagues would enjoy this newsletter, go ahead and forward this email to them. This will help us spread the word!
Semper discentes,
The D4S Team
A 14-billion-parameter model claims 28 gigabytes of memory before the first token. The KV cache then grows with every token of every open request, and real traffic opens many at once. This 374-page manual lives inside that squeeze. Chi Wang runs model-inference engineering teams at Salesforce. Peiheng Hu builds distributed inference engines at NVIDIA. The two spent over eight years building AI systems together, and it shows. Their claim is blunt. Training gets the papers, and serving carries the product. Welcome to the inference era.
The method sets Hands-On LLM Serving and Optimization apart. You build a serving service from scratch, batching and streaming included, and only then meet vLLM, TensorRT-LLM, SGLang, and llama.cpp. The frameworks stop looking like magic. The arithmetic alone earns the cover price. Estimate the weight footprint, size the KV cache, compute the arithmetic intensity of prefill and decode, and you know whether a workload is compute-bound or memory-bound before renting a single GPU. Continuous batching, quantization, speculative decoding, and the four parallelisms follow, each with runnable code, and an eight-step tuning plan for Qwen3-14B ties the whole book together. Data scientists get the systems course their training skipped. Platform engineers get a build-or-buy playbook.
Two warnings before you buy. The book went to press in April 2026, pinned to Qwen3-14B and today’s engine flags, and those engines ship faster than any press run. Treat the companion repository as the living copy. The second warning concerns fit. Many notebooks want an NVIDIA GPU, and the book stays on the serving side of the wall. Training, fine-tuning, and evals sit outside its scope, so readers hunting for modeling content will find queues and schedulers instead. Read it anyway. The publisher clocks it at 11 hours, one weekend. The serving layer decides what a model costs and how fast it feels, and most teams inherit it as a black box. This book opens the box and labels the parts. The next time the latency graph spikes or the GPU bill doubles, you will know which knob to turn first.
Opportunities to learn from us.