Industry

Everything · newest first

Models, products and AI startups, near-duplicates collapsed to the most credible source

GPT-5.6 Luna vs GPT-6 Astra for Code Review

In a 50-PR comparison, GPT-5.6 Luna found 69 verified bugs while GPT-6 Astra found 92; Luna run cost $0.20 versus Astra $5.66. Luna is much cheaper and handles routine correctness bugs well but has more false findings and finds fewer security issues, so it shouldn’t be sole reviewer for security-sensitive code.

Hacker News · · Details

Agents can't enable recursive self‑improvement

A new paper had authors hand agents accepted-but-unpublished NeurIPS papers (evaluated by the original authors); Codex/GPT-5.6 Sol and OpenClaw/Opus 4.8 failed to reproduce open-ended ML research. The authors argue this implies current agents cannot drive recursive self‑improvement (RSI), questioning near‑term explosive AI progress.

r/MachineLearning · · Details

Accountability Gap in Federal AI Use

Federal agencies are deploying AI in cybersecurity faster than governance can keep up. A Market Connections survey found 79% require human-in-the-loop for sensitive data, but under one-third have oversight frameworks and only 20% have predeployment testing policies, creating an accumulating accountability gap and risk.

FedTech Magazine · · Details

AI needs a full stop, not a slowdown

Parmy Olson argues in Bloomberg Opinion that calls to “pace” frontier AI are insufficient for existential risks, questioning Anthropic’s safety‑first stance amid competitive development and suggesting independent evaluations could improve oversight without necessarily slowing progress toward superintelligence.

Bloomberg · · Details

Can independent testing make AI safer?

Rayan Krishnan, co‑founder and CEO of Vals AI, says investment in independent testing lags model capability gains. Vals evaluates models from firms like OpenAI and Anthropic, sees early signs of recursive self‑improvement, but notes models remain well behind top human researchers.

Bloomberg · · Details

Pion: an agent to run companies autonomously

Andon released Pion, a platform to run real businesses autonomously (vending machines, stores, cafes) and opened a waitlist. They use Vending‑Bench to measure long‑term operation; Claude Opus 4 reportedly beat the human baseline in May 2025. Pion aims to study models’ real‑world resource acquisition and long‑term planning.

Hacker News · · Details

Will AI End Humanity? Deeper Questions Persist

An opinion piece links recent Silicon Valley warnings—including Anthropic CEO Dario Amodei's manifesto and a high-profile resignation—about AI-driven human extinction to older, historically similar technological fears (e.g., concerns during the Trinity nuclear test), urging a deeper look at the roots of such risk debates.

AlbertMohler.com · · Details

Why don't ML research agents overfit?

New research finds ML agents learn highly compressible strategies rather than memorizing data. Squeezing a successful agent’s strategy through an information bottleneck—down to about 16 tokens—still lets a fresh agent reproduce performance, showing compression distinguishes true generalization from overfitting.

Hacker News · · Details

When LLM judges agree, should we believe them?

The discussion proposes dependence-aware label aggregation using Ising models to model correlations among LLM judge panels, distinguishing independent evidence from shared mistakes; this improves accuracy by 9–14% in tests and recommends reporting confidence adjusted for judge correlation.

Hacker News · · Details