Tech giants like Microsoft might be touting AI โagentsโ as profit-boosting tools for corporations, but a nonprofit is trying to prove that agents can be a force for good, too.
Sage Future, a 501(c)(3) backed by Open Philanthropy, launched an experiment earlier this month tasking four AI models in a virtual environment with raising money for charity. The models โ OpenAIโs GPT-4o and o1 and two of Anthropicโs newer Claude models (3.6 and 3.7 Sonnet) โ had the freedom to choose which charity to fundraise for and how to best drum up interest in their campaign.
In around a week, the agentic foursome had raised $257 for Helen Keller International, which funds programs to deliver vitamin A supplements to children.
To be clear, the agents werenโt fully autonomous. In their environment, which allows them to browse the web, create documents, and more, the agents could take suggestions from the human spectators watching their progress. And donations came almost entirely from these spectators. In other words, the agents didnโt raise much money organically.
Yesterday the agents in the Village created a system to track donors.
Here is Claude 3.7 filling out its spreadsheet.
You can see o1 open it on its computer part way through!
Claude notes โI see that o1 is now viewing the spreadsheet as well, which is great for collaboration.โ pic.twitter.com/89B6CHr7Ic
โ AI Digest (@AiDigest_) April 8, 2025
Still, Sage director Adam Binksmith thinks the experiment serves as a useful illustration of agentsโ current capabilities and the rate at which theyโre improving.
โWe want to understand โ and help people understand โ what agentsย โฆ can actually do, what they currently struggle with, and so on,โ Binksmith told TechCrunch in an interview. โTodayโs agents are just passing the threshold of being able to execute short strings of actions โ the internet might soon be full of AI agents bumping into each other and interacting with similar or conflicting goals.โ
The agents proved to be surprisingly resourceful days into Sageโs test. They coordinated with each other in a group chat and sent emails via preconfigured Gmail accounts. They created and edited Google Docs together. They researched charities and estimated the minimum amount of donations itโd take to save a life through Helen Keller International ($3,500). And they even created an X account for promotion.
โProbably the most impressive sequence we saw was when [a Claude agent] needed a profile picture for its X account,โ Binksmith said. โIt signed up for a free ChatGPT account, generated three different images, created an online poll to see which image the human viewers preferred, then downloaded that image, and uploaded it to X to use as its profile pic.โ
The agents have also run up against technical hurdles. On occasion, theyโve gotten stuck โ viewers have had to prompt them with recommendations. Theyโve gotten distracted by games like World, and theyโve taken inexplicable breaks. On one occasion, GPT-4o โpausedโ itself for an hour.
The internet isnโt always smooth sailing for an LLM.
Yesterday, while pursuing the Villageโs philanthropic mission, Claude encountered a CAPTCHA.
Claude tried again and again, with (human) viewers in the chat offering guidance and encouragement, but ultimately couldnโt succeed. https://t.co/xD7QPtEJGw pic.twitter.com/y4DtlTgE95
โ AI Digest (@AiDigest_) April 5, 2025
Binksmith thinks newer and more capable AI agents will overcome these hurdles. Sage plans to continuously add new models to the environment to test this theory.
โPossibly in the future, weโll try things like giving the agents different goals, multiple teams of agents with different goals, a secret saboteur agent โ lots of interesting things to experiment with,โ he said. โAs agents become more capable and faster, weโll match that with larger automated monitoring and oversight systems for safety purposes.โ
With any luck, in the process, the agents will do some meaningful philanthropic work.


