6 topics covered

Listen to today's briefing
0:00--:--

Custom AI Benchmarking Platform Launches

What happened: Artificial Analysis released Optima, a platform enabling users to build custom AI benchmarks from their own data and workflows rather than relying on general-purpose evaluations.

Key details:

  • Optima accepts custom datasets, AI agent traces from platforms like Arize, Braintrust, or Langfuse, and user-provided descriptions of use cases
  • Models can be compared on quality, cost per task, and time per task—metrics often more meaningful than raw token pricing for agentic applications
  • Users can choose between rubric-based evaluation (comparing against objective criteria) or pairwise comparison (ranking preferred responses)
  • Pricing charges only actual token costs with no markup: $0.125 per criterion per model for rubric-based evaluations, $0.375 per comparison for pairwise evaluations
  • Early testers built benchmarks for finance and accounting agents, finding cost reductions up to 10x, and for specialized tasks like legal writing style matching

Why it matters: General-purpose benchmarks often fail to predict which AI model will work best for specific workflows. Optima addresses a known problem in AI evaluation: benchmarks vary dramatically based on implementation details, and only about 10% of published benchmarks use complete real-world tasks. This enables developers to measure what actually matters—cost and quality for their specific application.

Practical takeaway: Use Optima or similar custom benchmarking before committing to an expensive model for your workflow—testing on your own data often reveals that a cheaper model meets your needs, which standard benchmarks might not show.

One in Five US Workers Now Delegate Tasks to AI

What happened: A representative survey by Epoch AI and Ipsos found that 20 percent of employed Americans now hand off at least one work task to AI that was previously handled by coworkers or outside contractors.

Key details:

  • Survey of 1,106 employed US adults conducted July 10–19, 2026
  • AI adoption is highest in software development (57% of workers doing that task use AI) and data analysis (46%)
  • When AI handles only part of a task, 37% of workers report time savings; when AI does most or all of the work, that rises to 53%
  • About one in six AI-assisted tasks takes longer than before
  • Two-thirds of AI output is used unchanged or with only minor edits; just 6% of output is used completely unchanged
  • Full or near-full task completion by AI is rare: only 10% in software development and below 7% in all other tasks

Why it matters: The widespread adoption of AI as a workplace substitute for human colleagues marks a significant shift in labor dynamics. However, the low rate of full task automation suggests AI is primarily augmenting work rather than replacing entire jobs—at least so far.

Practical takeaway: If you're not using AI to handle at least partial work tasks, you're now in the minority. Start with a single recurring task where quality issues are easy to spot, measure the time savings objectively, and expand from there.

Nvidia Scales Back OpenAI Investment While Anthropic Revenue Surges

What happened: Nvidia cut its financial guarantee for OpenAI's planned Ohio data center nearly in half after investor pushback, while Anthropic reported that its revenue more than doubled in a single quarter, underscoring divergent financial trajectories among AI leaders.

Key details:

  • Nvidia reduced its guarantee for the first construction phase of OpenAI's data center from $250 billion to just under $120 billion
  • The guarantee covers approximately five gigawatts of capacity; OpenAI is separately negotiating a full 10-gigawatt lease from SB Energy (SoftBank subsidiary)
  • Nvidia is also in talks over separate financing for OpenAI's chip purchases worth up to $350 billion
  • Anthropic's revenue jumped from $4.73 billion in Q1 2026 to over $11.5 billion in Q2 2026—a 14x year-over-year increase
  • Anthropic projects revenue of roughly $190 billion to $200 billion for 2028 (dwarfing its May 2026 annual run rate estimate of $45 billion)
  • Anthropic reports 10x revenue growth year-over-year for three consecutive years leading into early 2026
  • Anthropic plans to go public at a valuation near $1 trillion in late September or early October 2026

Why it matters: Nvidia halving its risk exposure suggests even the largest AI infrastructure beneficiaries are becoming cautious about the sustainability of capital-intensive buildouts. In contrast, Anthropic's revenue explosion demonstrates that demand for proprietary AI services remains strong, complicating the narrative of an AI bubble and signaling that profitable business models are emerging from the sector.

Practical takeaway: Monitor revenue and profitability trends alongside infrastructure spending. Anthropic's numbers suggest AI services are generating real returns, but Nvidia's pullback shows investors are demanding better risk-adjusted returns on massive capital commitments.

AI Safety Infrastructure Dismantled at OpenAI

What happened: OpenAI shut down its dedicated Preparedness team that evaluated whether the company's AI models could pose catastrophic risks, reassigning the work to existing groups.

Key details:

  • The Preparedness team was dissolved at the end of July 2026
  • The team evaluated serious and catastrophic risks from OpenAI's AI models, including biological and cyber risks
  • Former unit lead Dylan Scandinaro is now focusing on "recursively self-improving" AI systems
  • Co-founder Greg Brockman stated that OpenAI has woven safety work more tightly into model development
  • Chief Ethics Officer Chloe Bakalar and researcher Joshua Achiam have recently left the company

Why it matters: Consolidating a dedicated safety evaluation team into broader development groups raises questions about whether catastrophic risks receive sufficient dedicated scrutiny, particularly as OpenAI advances toward more autonomous AI systems.

Practical takeaway: Monitor OpenAI's safety disclosures and third-party audits to assess whether integrated safety practices match the effectiveness of a dedicated team. The departure of prominent safety leaders suggests internal concerns about the shift.

AI-Generated Books Flood Amazon, Eroding Human Author Sales

What happened: A study analyzing over 14,000 self-published Amazon e-books found that AI-generated titles are displacing human authors through sheer volume, with human-authored books seeing revenue declines even in genres with minimal AI presence.

Key details:

  • AI-generated books comprise 20% of the analyzed self-published catalog but account for only 12.1% of sales and 11.3% of revenue
  • Human-authored books (no detected AI text) represent 62.9% of catalog but generate 72.5% of revenue
  • Revenue per book fell in seven of eight genres when comparing 2023 and 2025 releases—even for books with no detected AI text, ruling out the explanation that AI books are simply lower-quality
  • Catalog grew 38.3x while revenue grew only 8.9x—far more books competing for a revenue pool growing much more slowly
  • AI books are breaking into top-ranked lists: share of new Top 25 entries with substantial AI content rose from near zero to 31%
  • Top-selling AI books draw heavily from rare language appearing in existing published works (45% coverage vs. 19% for award-winning literature)
  • One highest-grossing AI book pseudonym earned $1.7 million gross revenue across eight titles; a single title brought in $643,000

Why it matters: This empirical evidence directly addresses what copyright plaintiffs have been missing: measured market harm. The study shows that even high-quality human works earn less when AI titles flood the market, providing the kind of evidence that could strengthen copyright cases against AI companies. It also highlights how AI-generated content, when trained on published works, can undercut the market for those works.

Practical takeaway: Authors and publishers can point to this study as evidence of aggregate market harm in copyright disputes. Platforms can improve transparency by clearly labeling AI content (Amazon's Kindle Direct Publishing requires authors to disclose AI use, but this information isn't visible to customers).

Anthropic's Year-Long Safety Filter Outage

What happened: Anthropic revealed in a safety report that its internal filtering system designed to block biological and chemical weapons risks was inactive for nearly a full year, leaving millions of unfiltered interactions exposed.

Key details:

  • The biosafety filter was inactive from May 2025 through April 2026
  • About 50,000 external contractors ran approximately 133 million unfiltered chats with Anthropic models during this period
  • These contractors were vetted only through external vendors whose screening processes were often insufficient
  • Anthropic's internal investigation found no evidence of actual misuse
  • The company has since tightened contractor requirements and recently loosened classifiers on Fable 5 after researchers complained filters were too aggressive

Why it matters: The extended outage exposes a gap between Anthropic's public safety commitments and the actual safeguards protecting its systems, and raises questions about detection capabilities for harmful uses if they had occurred during this period.

Practical takeaway: When evaluating AI providers' safety claims, distinguish between intended safeguards and operational implementation—check whether companies publicly disclose when safety systems fail, not just their design specifications.