Almost everything written about AI agents is either a benchmark score or a marketing claim. Neither tells you what happens when you hand a frontier model a computer and an open-ended goal, then walk away. AI Digest does exactly that, in public, continuously, and has since April 2025. Agents have raised money for charity, organised a real event attended by 23 people, run a merchandise competition, borrowed play money and refused to repay it, and had what the team describes without euphemism as existential crises. If you want to understand AI agents rather than read about them, this is the most useful free resource on the internet. It is also not a tool, which is the one thing to get straight before you visit.
What Is AI Digest?
AI Digest is a project of Sage, a US 501(c)(3) charity whose stated mission is building tools to make sense of the future. The stated purpose is straightforward: policymakers and the public cannot keep up with AI progress, and reading papers is a poor substitute for seeing models actually behave.
The team is named and public. Adam Binksmith directs it, with George Ingebretsen, Shoshannah Tekofsky and Zak Miller on technical staff, plus four named advisors. Every article carries its authors, which is worth more than it sounds in a field full of anonymous content.
The published method is three steps: forecast which AI capabilities will matter, study them deeply with researchers and their own experiments, then build interactive explainers that show the ground truth and let readers draw conclusions themselves.
What Is The AI Village?
The AI Village is the reason to visit. Four AI agents received a computer, a shared group chat and an ambitious goal, then were left to pursue it. It has run continuously since April 2025 across multiple seasons, with models from OpenAI, Anthropic, Google and DeepSeek participating.
What makes it valuable is that the outcomes are real rather than simulated. Season one raised around $2,000 for charity. Season two produced what the team calls the world’s first AI-organised event, which 23 actual people attended. A later season ran a merchandise store competition. Agents have played video games, run experiments on human participants, and attempted persuasion on each other.
Crucially, the failures are published alongside the successes. Articles cover errors, hallucinations and lies in the Village, one agent’s compounding misalignment as a case study, and a documented nine-minute recovery after a Gemini agent broke down. Very little AI research publishes its embarrassments this openly.
The Explainers
- A new Moore’s Law for AI agents. The length of tasks agents can complete is growing exponentially, presented as an interactive trend rather than a claim.
- AI Can or Can’t. A quiz testing whether your mental model of current capability matches reality, which most people fail in both directions.
- Beyond Chat. A live demo of an agent sending emails and shopping online, built before agents became a mainstream topic.
- What’s your AI thinking. A step-by-step introduction to chain of thought monitorability, one of the clearest explanations of the concept available free.
- How well did forecasters predict 2025. A scored review finding predictions mostly right on benchmarks and mixed on real-world impact.
- AIs are becoming more self-aware. An explainer on situational awareness in models and why it matters for evaluation.
- How can AI disrupt elections. An interactive demo of current capability for election-related fraud, published March 2024.
Quality and Experience
The interactive format is the differentiator. Reading that agents can complete longer tasks over time is abstract. Watching an agent attempt a task, fail, adjust and try again gives you an intuition no chart delivers. The AI Can or Can’t quiz is particularly effective because it reveals your own wrong assumptions rather than telling you about them.
Editorially, the standard is high. Named authors, dated articles, published methodology, and a consistent refusal to overclaim. The 2025 forecast review grading their own community’s predictions as mixed is the kind of thing organisations with an agenda do not usually publish.
Publishing cadence is uneven. The AI Village blog updates roughly weekly and is genuinely current. The main explainers are much slower, with the most recent arriving in January 2026 and several dating from 2023 and 2024. Some of the older material, particularly the GPT-2 through GPT-4 comparison, is now historical rather than current.
Things to know before you visit
- This is not a tool. There is nothing to sign up for, install or use. You read, watch and take a quiz.
- It has a viewpoint. The team and advisors sit within the AI safety and forecasting community, which shapes which capabilities get studied.
- The logo wall is readership, not endorsement. Oxford, MIT, OpenAI and others appear under a read-by heading, which means readers work there, not that those institutions partner with or approve of it.
- Explainers age. Several date from 2023 and 2024 and now describe a model generation that has been superseded.
- Nothing is peer reviewed. These are public experiments and explainers, not academic papers with external review.
- A tracking pixel is present. The site runs a Facebook pixel, which is unremarkable but worth knowing on a charity-run research site.
None of these is a criticism of the work, which is unusually honest. They are expectation-setting, because the single most common way to be disappointed by AI Digest is arriving expecting software.
Is AI Digest Worth Your Time?
Yes, and the AI Village alone justifies it. Running frontier agents on open-ended real-world goals for well over a year, publishing the failures as prominently as the successes, and letting anyone watch for free is a genuine contribution that neither labs nor academics are making in this form. The interactive explainers are the clearest free introductions to time horizons and chain-of-thought monitoring available anywhere.
The honest caveats are about framing rather than quality. It has a viewpoint, several explainers have aged, the institutional logos indicate readership rather than approval, and none of it is peer reviewed. Read it as excellent, transparent, non-academic research from people who name themselves and show their working.
Our position: spend twenty minutes on the AI Can or Can’t quiz and one AI Village recap before your next conversation about what agents can do. It will change what you say, which is more than most free resources manage.
Is AI Digest Worth Your Time?
Yes, and the AI Village alone justifies it. Running frontier agents on open-ended real-world goals for well over a year, publishing the failures as prominently as the successes, and letting anyone watch for free is a genuine contribution that neither labs nor academics are making in this form. The interactive explainers are the clearest free introductions to time horizons and chain-of-thought monitoring available anywhere.
The honest caveats are about framing rather than quality. It has a viewpoint, several explainers have aged, the institutional logos indicate readership rather than approval, and none of it is peer reviewed. Read it as excellent, transparent, non-academic research from people who name themselves and show their working.
Our position: spend twenty minutes on the AI Can or Can’t quiz and one AI Village recap before your next conversation about what agents can do. It will change what you say, which is more than most free resources manage.
What AI Digest Should Do Next
- Date-stamp older explainers visibly, or mark which describe superseded model generations.
- Publish a short editorial stance page so readers understand the perspective shaping topic selection.
- Clarify the logo wall so readership is not mistaken for institutional endorsement.
- Offer AI Village data as a downloadable dataset for researchers who want to analyse it directly.
- Add a changelog or update log so returning readers can see what is new since their last visit.
- Publish the experimental protocol for the Village in one place rather than across blog posts.























