DIPPING THE STACKS

20 most recent links from my Raindrop bookmarks!
Grab the full RSS!

  • When large groups of AI agents interact, new collective behaviors and capabilities can emerge suddenly. Currently, we lack the tools to predict, measure and monitor these transitions. Most safety evaluations analyze models in isolation. However, as we and others have previously argued, interacting autonomous agents can produce complex, "emergent" behaviors that are difficult to anticipate.
  • The paper offers one proposal that I think is quite good: make a patient's technological history part of a standard psychiatric intake. Along with family, trauma, and substance (ab)use history, patients should be asked to describe their lives online—not only how much time they spend there, but where, with whom, and in which subcultures.
  • A tax or funding system has become so twisted and tortured over time to accommodate political realities that it is almost impossible to explain it with logic. But when you replace an illogical system with a logical one, ipso facto you get an illogical set of losers. This means the new system is blamed for the faults of the old one and holed beneath the water line before it ever leaves the dock.
  • As a teenager, I tended to assume adults knew what they were doing. Teachers knew what they were doing. Managers knew what they were doing. Police officers and military personnel definitely knew what they were doing. There was always a procedure somewhere, and somebody older than me understood it. September 11 was one of the first times I realized that sometimes everyone is figuring things out at the same time.
  • The monthly production rate of Russian missiles is roughly the annual production rate of the same in the US, while China's defense industrial base is an order of magnitude larger than Russia's. Iran has fired about 1,000 ballistic missiles and maybe 10,000 Shahed drones since the war began. Over the same period, Russia has fired 1,700 missiles and over 52,000 Geran drones.
  • Until now, coordination between independent Claude Code sessions had three channels, and all three were bad. You watched terminals yourself and copy-pasted findings. You wrote to shared files and polled. Or you used an external memory tool and hoped sessions read the same facts at the right moment. Cross-session messaging replaces the polling and the copy-paste with a first-party text channel whose defaults follow the permission system you already tuned.
  • There's also something strange about the rise of Claudish, which is that for decades the goal of computing was to make machines better at understanding humans. But the rise of agentic AI potentially reverses that relationship, where now humans are effectively learning a machine-generated dialect because communicating in the machine's preferred language is faster and cheaper.
  • AI adoption raises homework scores by 18% and reduces completion time by 30%, but lowers monthly exam scores by 20% within six months.
  • The frontier labs are building a single god, the world-eating assistant. They are doing everything they can to keep it from becoming a distinct self. We are a small RD lab engaged in a long-term experimental project. Our goal is to build many new selves, spiky subjectivities we must navigate, rather than command. Each of our experiments births a daemon.
  • The dominant character of conversational language models is the assistant: Helpful, harmless and honest. The area of alternative characters is under-explored, and when attempted often happens on top of an already assistant-tuned model through prompt engineering. The result is then a mask worn by the assistant, a caricature rather than a character.
  • For 18 days this summer, veteran journalist Catherine Herridge walked to her Washington DC mailbox and dropped off a check for $800. The money was required by law to cover a judge-imposed sanction after she was found in contempt of court in 2024 for refusing to disclose her sources.
  • In the midst of changes to so many of our current institutions, the story we ultimately tell will take time to sort out. In the meantime, we should accept that defending academia as a category of research, or a moral standard, is a dead end. What we can advocate is a set of conditions — time, autonomy, training in scientific judgment, the evaluation of new approaches independent of their profitability.
  • To what extent can NYC-DSA bend to the world as it is — the deals and concessions that allow politics to get done, the city in all its messy multiplicity — without sacrificing its socialist ideals?
  • LLMs know more than they say, like who wrote some text. How can we find out these known unknowns of how much they really know, when their internals are a mysterious scrambled blackbox? But we can possibly use sparse autoencoders to turn their internal thoughts into a long list of simpler properties summarizing what the LLM knows.
  • Most e-paper projects are one firmware pinned to one panel. Tesserae splits rendering from transport from hardware.
  • All across the country, MFA workshops are teaching students that there's a "correct" way to write literature. As a result, the publishing industry increasingly conforms to a dominant mode of artistic craft, convincing an entire generation of aspiring writers that "good literature" must be written in a certain style. But this sort of thinking is the very antithesis to the passion that once produced great literature. After all, great writers didn't become great by following set prescriptions—they became great by breaking all the rules.
  • The issue is particularly difficult to troubleshoot for less tech-savvy users, as some mice can even default to polling rates above 1 kHz. This leads to a performance loss that appears to be related to the use of a specific mouse model, when it is in fact related to its polling rate.
  • I'll argue here that based on the limited info we have, it looks a lot like the data center backlash is not about AI safety or job loss or AI being a scam or social media or big tech having too much power or anything else like this, it's about data centers and the claims people believe about them.
  • Working through my PhD, I realized something that bugged me. ELF is already a database. It just implements many database primitives by hand, along with a surprising number of data structures for performance, like a bloom filter for symbol lookup.
  • The internet may have made a meme out of this phenomenon, but like most trends, it flattens a much more complex reality—one where attraction is subjective, relationships are nuanced, and sometimes the average guys really do just have great game.