It’s rare to find an investment that pays off immediately and astronomically, but a treasure hunter scouring the Michigan estate sale of a local physician scored the jackpot: he bought a painting of several shells for $30 that was ultimately discovered to be “Alderman Merriam’s Shells,” a 1951 painting by Gertrude Abercrombie. The artist had traded the painting to the doctor in exchange for medical services. The painting just sold at auction for $1.3 million.
Stocks fell on Wednesday.
Last week, Google’s DeepMind announced that it was releasing Gemini Argon 4 to a select group of testers, but it was the specs of the model and its performance on a battery of exams that really got people’s attention.
That performance also has some significant relevance to investors, because if the benchmarks hold up, Google has managed to succeed in a number of areas where its competitors have ceded ground:
While the major AI labs have all been churning out products on the frontier, amid all the chaos of new models — Opus 5.5! Fable! Astra! — the understanding of who commands the bleeding edge of AI development has actually been rather boring.
Based on LMArena, it’s been Anthropic’s Claude dominating the rankings, and since March Claude has been a strong favorite on prediction markets to finish the year as the best AI.
That’s why Argon is so interesting: even though it hasn’t been rolled out to customers, the performance of the model was so solid that, in the week since it was revealed, Gemini jumped neck-and-neck with Claude*.
Now, let’s talk those metrics.
The details on Argon’s benchmarking posted by Demis Hassabis, Alphabet’s top AI honcho, made it clear that it’s good at navigating dense legalese, is top of the line at business automation, and makes stuff up way less:
Argon appears to have crushed the competition on Harvey’s Legal Agent Benchmark, coming in at a score of 19.6% compared to nearest runner-up Claude Fable’s 6.7%.
Also of note is the performance on AutomationBench, a benchmark from Zapier about how good models are at basic daily tasks in sales (such as lead management, CRM), marketing (ad performance), finance (expenses and bookkeeping) and HR (payroll and onboarding). There, it outran the nearest competition by nearly ten points.
Those analyzing the model’s performance have also found something else impressive: Argon also appears to have knocked it out of the park when it comes to reining in hallucinations. According to the AA-Omniscience Hallucination Rate, which calculates how often a model answers incorrectly when it gives a partial or non-response, is the lowest on the frontier, at just 15%. That’s obviously not zero, but a major improvement compared to GPT-6 Astra (45%) and Opus 5.5 (59%).
This doesn’t mean you can replace your lawyer with Argon, but… if you are a company that earns a solid slice of billable hours consulting, offering basic legal advice, human resource services, or actionable intelligence, I might be a little worried about the AI in every client’s email getting sharper at precisely these tasks.
The Takeaway
These developments might be especially relevant for Google’s business in particular. Several of their peers in AI development have understandably focused their efforts on designing models that are particularly good at coding.
From a long-term ROI perspective, it makes sense for technology companies to invest in tools that can build better tools. There’s also tons of online material about how to code for AI to ingest. And of course, humans who develop software are just incredibly expensive, and if you can build a robot that develops software you can sell it to a lot of fabulously wealthy technology companies and make a lot of money before you IPO.
While some of those incentives are advantageous for Alphabet, the more important element here is to think about what makes Alphabet’s core customer more unique. Nevermind coders, Alphabet has lots and lots of corporate G Suite customers and even more people who just use elements of the Google family of products.
How Google will make money from AI is by creating tools that can be used by the companies and people that already pay Google. In other words, Alphabet’s best shot at dominating an AI niche is to build within the business services niche it already dominates, selling to the hundreds of millions of people who work non-software jobs who want useful tools for those jobs.
That’s why the Argon result is just so eye-popping, and why the rollout of Argon could be a make-or-break moment not just for Alphabet, but for the scores of companies that Argon may very well disrupt. Just ask the software developers and SaaS companies how fun it was for them when Claude would demonstrate on a weekly basis how delicious it found their lunch to be.
Glenn Close clips. 1950s commercials. Uzbekistan’s version of ‘America’s Got Talent.’ The legendary investor can’t resist the algorithm either.
🏈 NFL: We’re a few weeks in, and the teams expected to appear in the Super Bowl have not changed all that much. The Rams remain a 21% chance to represent the NFC, followed by the 49ers (18%), Seahawks (16%) and, in a bit of a surprise, the Bears (11%). From the AFC the leading contender is the Bills (20%), followed by the Chiefs (19%), Ravens (18%) and Jaguars (13%)
🎵 Halftime: The only thing less clear than the teams on the field is the performer at halftime, but there are a few leading contenders for the big show, with markets putting it at a 47% chance of JAY-Z performing, a 34% chance of Miley Cyrus, 21% chance of Harry Styles, and 11% chance of Justin Bieber.
*Event contracts are offered through Robinhood Derivatives, LLC — probabilities referenced or sourced from KalshiEx LLC or ForecastEx LLC.
That’s the average drive-thru time in America, according to the annual QSR Magazine report. Speed isn’t everything, though: Chick-Fil-A scored some of the best ratings for a drive thru even though their experience averaged 7 minutes, 5 seconds.
They paid $50,000 and waited nine years. Will Tesla’s Roadster live up to the hype?
Apollo, banks in talks to finance SpaceX’s $40 billion Nvidia GPU purchase
Disney+ to stream upcoming Super Bowl
Global bond sell-off resumes as 30-year Treasury yield hits highest since 2002
Beyond omakase, sushi bars are hosting 200-pound tuna carving parties
An AI has finally cracked Stratego, one of the hardest games for a computer to master because it has 40 pieces per side, a decillion possible setups, and imperfect information. Ataraxos just went 15-1-4 against the world’s best player.
PepsiCo slated to release quarterly results ahead of the open.
St Louis Fed President Alberto Musalem due to speak at 1:40 p.m. ET.