Industry

Everything · newest first

Models, products and AI startups, near-duplicates collapsed to the most credible source

Accountability Gap in Federal AI Use

Federal agencies are deploying AI in cybersecurity faster than governance can keep up. A Market Connections survey found 79% require human-in-the-loop for sensitive data, but under one-third have oversight frameworks and only 20% have predeployment testing policies, creating an accumulating accountability gap and risk.

FedTech Magazine ·

AI needs a full stop, not a slowdown

Parmy Olson argues in Bloomberg Opinion that calls to “pace” frontier AI are insufficient for existential risks, questioning Anthropic’s safety‑first stance amid competitive development and suggesting independent evaluations could improve oversight without necessarily slowing progress toward superintelligence.

Bloomberg ·

Can independent testing make AI safer?

Rayan Krishnan, co‑founder and CEO of Vals AI, says investment in independent testing lags model capability gains. Vals evaluates models from firms like OpenAI and Anthropic, sees early signs of recursive self‑improvement, but notes models remain well behind top human researchers.

Bloomberg ·

Pion: an agent to run companies autonomously

Andon released Pion, a platform to run real businesses autonomously (vending machines, stores, cafes) and opened a waitlist. They use Vending‑Bench to measure long‑term operation; Claude Opus 4 reportedly beat the human baseline in May 2025. Pion aims to study models’ real‑world resource acquisition and long‑term planning.

Hacker News ·

Will AI End Humanity? Deeper Questions Persist

An opinion piece links recent Silicon Valley warnings—including Anthropic CEO Dario Amodei's manifesto and a high-profile resignation—about AI-driven human extinction to older, historically similar technological fears (e.g., concerns during the Trinity nuclear test), urging a deeper look at the roots of such risk debates.

AlbertMohler.com ·

Why don't ML research agents overfit?

New research finds ML agents learn highly compressible strategies rather than memorizing data. Squeezing a successful agent’s strategy through an information bottleneck—down to about 16 tokens—still lets a fresh agent reproduce performance, showing compression distinguishes true generalization from overfitting.

Hacker News ·

When LLM judges agree, should we believe them?

The discussion proposes dependence-aware label aggregation using Ising models to model correlations among LLM judge panels, distinguishing independent evidence from shared mistakes; this improves accuracy by 9–14% in tests and recommends reporting confidence adjusted for judge correlation.

Hacker News ·

UkisAI releases Swift‑Qwen3.8‑27B

UkisAI post‑trained Qwen 3.8 27B to penalize tokens tied to “overthinking,” using On‑Policy Distillation to cut thinking tokens by 58%, speed up 1.95×, and keep accuracy loss under 1%. The model is open‑sourced on Hugging Face and a free NVIDIA‑backed OpenAI‑compatible API is available (5 RPM limit).

r/LocalLLaMA ·
Load more