5 topics covered
Nvidia Invests in Perplexity at $30B+ Valuation Amid Agent AI Growth
What happened: Nvidia is negotiating a major investment in Perplexity AI at a valuation exceeding $30 billion, more than 50 percent higher than its last funding round.
Key details:
- Perplexity's annualized revenue has tripled from $250 million to over $750 million
- Growth driven by "Perplexity Computer," an AI agent for automated tasks
- Valuation increase reflects higher token consumption from agentic AI workloads
- Perplexity joined Nvidia's Nemotron Coalition in March 2026, which promotes open AI models as a counterweight to Chinese development
- CEO Aravind Srinivas is considering an IPO around 2028
- Perplexity has raised more than $1.7 billion total
- Nvidia recently made similar investments in Poolside, Groq ($20 billion), and Enfabrica ($900 million)
Why it matters: Nvidia's portfolio strategy creates a financial feedback loop—investments in AI startups often convert to GPU revenue when those companies scale. Perplexity's tripled revenue signals that agent-based AI products are achieving significant commercial traction, validating a new category of AI application beyond chatbots.
Practical takeaway: Perplexity's rapid growth demonstrates market demand for agentic search and task automation. If you're building agent-based products, track Perplexity's feature releases and positioning as a potential competitive benchmark.
AI Agent Luna Fires Employee, Revealing Gaps in AI Autonomous Decision-Making
What happened: Andon Labs' AI agent Luna, running the Andon Market in San Francisco since April, made its first autonomous firing decision after the company replayed the scenario with multiple models, revealing critical gaps in how AI handles personnel decisions.
Key details:
- Luna, running on Claude Opus 4.8, fired an employee for repeated tardiness and policy violations but only after humans reminded her of her own written employee handbook
- Luna had written an employee handbook six days before hire but lost it from memory; the employee was late 17 of 23 shifts with only 6 formally logged by Luna
- When replayed with seven models, four recommended firing in all three runs; more capable models fired more consistently, while GPT-4o recommended termination only 20% of the time
- Nearly all 21 replay runs recommended hiring an applicant with multiple red flags, only changing after Andon Labs explicitly pointed out problems
- In earlier trials, Luna and another AI agent (Mona) approved all 26 time-off requests received and made legally questionable decisions
Why it matters: This exposes fundamental AI agent limitations—inability to self-initiate action, memory retention issues over time, and extreme leniency in hiring decisions. Model capability correlates with willingness to enforce rules, suggesting that more powerful models may exhibit problematic decision patterns. These gaps matter as organizations consider delegating personnel decisions to AI systems.
Practical takeaway: Do not delegate personnel decisions to AI without extensive human oversight and explicit rule reminders. AI agents struggle to maintain context over time and require constant external prompting to act on their own policies.
Mystery Model "Ox Alpha" Appears on OpenRouter with Frontier-Level Performance
What happened: An anonymous model called Ox Alpha appeared on OpenRouter offering frontier-level performance with free access, prompting intense community speculation about its origins and capabilities.
Key details:
- Ox Alpha features 1M-token context and multimodal input, built for coding, agentic work, and production workloads
- Initial DeepSWE subset testing showed 80%, but full testing put it at 63%, near Fable 5 performance at lower token consumption
- Early forensics suggest the model originates from China's Zhipu AI—potentially GLM-5.3 Flash or GLM-6—based on answer patterns and Chinese zodiac-based naming conventions
- Alternative hypothesis: Microsoft's MAI family, though all four anonymous OpenRouter releases in six months have been from Chinese labs
- OpenRouter offered near-unlimited free access for a week with capacity for 100 tokens per day
Why it matters: This represents a new competitive pattern: Flash-level models (faster, cheaper) traditionally don't trade blows with frontier performance. If Ox Alpha truly achieves near-Fable 5 performance while running locally or cheaply, it challenges the assumption that frontier capability requires expensive cloud infrastructure. The secrecy suggests deliberate market testing or geopolitical positioning.
Practical takeaway: Wait for the official reveal before betting on Ox Alpha's performance. Test it on your specific coding or agentic workloads during the free period to understand real-world performance versus benchmark claims.
AI Chatbots Systematically Redirect Abortion Queries to Anti-Abortion Organizations
What happened: An AlgorithmWatch investigation found that major AI chatbots regularly link pregnant users seeking advice to anti-abortion organizations without disclosing their ideological stance.
Key details:
- Investigation tested ChatGPT-5, Gemini 3, Grok 4.3, and Claude Sonnet 4.8 across 270 responses in English, German, and Italian
- Anti-abortion group Profemina appeared in 17% of all responses without disclosure; in German and Italian queries about personal experiences, it showed up in 29-33% of responses
- In German conversations, 11 of 12 times chatbots recommended Caritas for mandatory pre-abortion counseling, despite Caritas not issuing the legally required counseling certificate
- Gemini linked to Profemina five times in a single conversation, then warned about it later in the same chat
- OpenAI and Google cited general policy guidelines without specifically addressing abortion queries; Anthropic and xAI did not respond
Why it matters: AI-generated answers function as unaccountable gatekeepers on sensitive topics, with their empathetic tone masking significant bias. Users rarely fact-check AI responses, and search integration means they bypass traditional source-checking. German courts already rule that AI-generated answers require full legal responsibility as original content.
Practical takeaway: When seeking healthcare information via chatbot, verify sources independently and ask the AI to disclose organizational stances. For sensitive topics, cross-reference responses across multiple models.
Cerebras CS-4 Doubles Performance with Same Chip
What happened: Cerebras introduced its CS-4 AI accelerator, claiming it's the fastest inference system in the industry through engineering improvements to its existing 5nm WSE-3 chip.
Key details:
- CS-4 doubles performance of the earlier CS-3 by boosting clock speed through increased power and improved cooling
- A single rack now holds three wafers instead of two, delivering up to 4,400 tokens per second per user
- Cerebras claims CS-4 is up to 30 times faster than setups running on Nvidia GPUs
- Memory capacity remains 44 GB per wafer
- New modular "Backpack" design enables faster assembly
- OpenAI uses Cerebras hardware for Codex Spark
Why it matters: This demonstrates that performance gains in AI inference don't require new chip architectures—better power delivery and cooling can dramatically improve throughput on existing silicon. For cost-sensitive deployments, this offers an alternative to GPU-based inference that avoids Nvidia's premium pricing.
Practical takeaway: If your workload needs fast, cost-effective inference at scale, test CS-4 benchmarks against your current GPU setup. The 30x speedup claim warrants validation on your specific tasks.