We’ve had an AI breakthrough at Figure and will be showcasing this tomorrow
What leading AI figures posted on X
Brett Adcock said Figure achieved an AI breakthrough and will showcase it tomorrow.
Nathan Lambert praised the public MiMo-V2.6 large-scale RL resource as very impressive.
Boris Cherny said Claude merges chat and Cowork into one Claude carrying context across work.
Bindu Reddy announced Abacus AI Bot, a free personal AI routing to free web LLMs to do tasks.
Alexandr Wang said Muse saved users $9,649.71 across 100 real stories.
Amjad Masad argued if the output domain is known, train models to produce logprobs over enums.
We’ve had an AI breakthrough at Figure and will be showcasing this tomorrow
excited to help Shopify merchants advertise their products in ChatGPT: [Quoting @harleyf]: ChatGPT Ads for @Shopify is live. We are @OpenAI's first commerce partner. OpenAI pulls straight from Shopify Catalog, so merchants' products are already there. Mer…
wall-to-wall deployment of astra for engineers at databricks: [Quoting @pwendell]: Today we rolled out Astra to every engineer at Databricks (N=~3500). Some notes that may be helpful to others: 1. Astra unambiguously out performs our previous highest-end mode…
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. Mo…
One of the coolest at-scale RL resources made public yet! You love to see it. [Quoting @_LuoFuli]: Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scal…
Use Grok Voice in fal to build intelligent, low-latency agents that resolve real customer issues [Quoting @fal]: Grok Voice is live on fal. The latest speech model from @SpaceXAI that answers in 0.70 seconds and finishes its tool calls before the sentence end…
We're seeing extraordinary results from @typesafeai. Default mode in 𝚏𝚡 is auto, with a safety reviewer analyzing every command. That reviewer runs on GPT Luna today. Jev is up to 18x faster (p95) *and* more accurate. It's coming to @vercel AI Gateway and l…
We rolled out GPT-6 Astra to every Databricks engineer today. It beats Claude Opus 5 on the hardest, long-horizon tasks and increased our coding spend by 60%. I think it’s the strongest model. Expensive. But hopefully worth it. [Quoting @pwendell]: Today we r…
Grok Build now gets better the more you use it. It remembers conventions, decisions, and project facts across sessions. /memory to browse /dream to organize recent notes into topics https://t.co/A2MMZT5eJX https://t.co/UD8lJGnfOB
Super interesting read! [Quoting @JacquelineSYC19]: A few weeks ago I wrote that our whole team at Artie uses Hermes (an AI agent harness from @NousResearch) to do their work. Here's how we got there: what we tried first, what it cost us, how we built a Herm…
"muse is the best app - AI or otherwise - I've used since ChatGPT in terms of bringing back that original 'wow' factor." you're making me emotional 🥹🥹🥹 [Quoting @0xShayan]: Muse is the best app - AI or otherwise - I've used since ChatGPT in terms of bring…
You can just build (physical) things. [Quoting @natalieyeo]: "Hello World" from this cute physical Codex pet! To all the @OpenAIDevs Day 5 of building with Astra https://t.co/aTeijhJQT1
Gave my mom my (very simple) agent setup. She named him Morgan. He just closed a $600 ad deal for her site. People are massively overthinking their setups. Bitter Lesson applies to agents too: just ask the model, let it spawn what it needs. https://t.co/ldJzE…
Codex voice for road trips: [Quoting @jonathanroomer]: I’ve got Codex (voice) in CarPlay. And it’s fantastic. I can now build things during road trips, while all 5 kids make a noise and my wife asks why I talk to the AI more than I talk to her. It runs throug…
Introducing the Hermes Agent plugins catalog! You can access it in your Hermes Desktop app's capabilities section to easily search, browse, and explore community and official plugins! Submit your own as well! [Quoting @NousResearch]: Hermes Agent now has a Pl…
Any guess what this Union Alpha beast on OpenRouter is? Claude Fable 5.1-level, over 10x cheaper. Waiting for it to be open-sourced. https://t.co/okoQnHeMw6
Also today: Claude Docs, Claude Slides, and Claude Design are in every conversation. Ask Claude for a presentation and you get one you can open, edit, and export as PowerPoint or PDF. Same for a document or a design. There's no separate tool to navigate to. T…
Claude Code showed that AI could do real work, not just answer questions. Developers hand Claude a feature, come back to shipped code. That's where much of the industry's serious engineering runs now. Cowork proved knowledge workers could do the same: hand Cl…
Three key open-weight labs have seen monthly dollars spent on their models grow by 10x or more so far in 2026. The closed models have seen comparatively slower growth in spending, though from much higher starting points. https://t.co/iOKARyF2as
When I talk to senior managers, increasingly hearing stories of the blurring of jobs inside organizations (we found this at our P&G study as well): coding, design, product management, all collapsing & overlapping as everyone uses AI We need new models…
muse voice transcribe is really good!! [Quoting @IsaacKing314]: I regret to inform you all that Meta AI is finally good, at least in the fields where their competitors have stopped trying. Their new voice transcription model is nearly an order of magnitude fa…
Haha... Literally no one is pacing the frontier - Opus 5.2 in testing - Grok 4.8 ships in a couple of weeks - Jev is a new ultra fast classifier - OpenAI already has Astra+ in testing We continue to accelerate
saying weird stuff about AI and the singularity The Anthropic Institute 🤝 the DeepMind Institute [Quoting @ShaneLegg]: My journey to develop AGI spans 25 yrs, including 10+ yrs thinking about technical & societal perspectives at Google DeepMind. AGI is…
Thrilled to complete the definitive and join forces with AlephAlpha! 🇨🇦🇩🇪 [Quoting @cohere]: Cohere and Aleph Alpha announce the signing of a definitive agreement, becoming the first foundational AI model developer anchored on both sides of the Atlantic …
🚨 Announcing Abacus AI Bot - 100% FREE PERSONAL AI AI can manage your calendar, book tickets, or make reservations. These simple tasks should be TOTALLY FREE Today, we announce our new product - Abacus AI bot - Finds and runs on FREE LLMs on the web - Autom…
Today, we are announcing a partnership with @mozilla to bring privacy, control and choice to people using AI to browse online. 🦊🐈 https://t.co/Bsd6N4jNdV https://t.co/OczDk9utfL
Study from Google about AI use in science, complex impacts: acceleration (7 hours saved per week) along with shifts in the kind of work (more verification) and what research gets done (possibly safer topics). Also a good diagram of the jagged frontier https:/…
muse saved this person over $5000 on WiFi! [Quoting @raunaqbn]: @Muse just continues to blow my mind! Today I had it call Xfinity to haggle down my internet bill. it got through the phone tree to a human, hit the verification text it couldn't read, and patche…
muse saved people $9,649.71 across 100 different stories imagine how much money this’ll save people when a billion people have muse!! we are working hard to expand access and get it everywhere! [Quoting @SingularityRes]: Meta's Muse AI has saved real people $…
i swear muse code is actually good!! the same model behind muse app powers muse code! [Quoting @ryanmcadams]: Never thought I'd say this, but @alexandr_wang has been telling people how good Muse code is, and I wasn't buying it. Spent real time in it today and…
Jev has spoken. It picked which model is AGI. 20–200x faster. 40–400x cheaper. This could make things like LLM-as-a-judge insanely fast and nearly free. (I tried a bunch of prompts and still didn’t burn through $0.10.) https://t.co/1Uhh4nMGDl [Quoting @Comple…
Zuck just swooped in and solved AI pacing! - labs should make models safe - If they don’t they will face criminal liability when the AI goes rogue - we have the gaurdrails in place already and Meta already paced their AI and made sure it’s safe Problem solv…
This is cool, but if your output domain is known in advance, why not just train a model to produce logprobs over enums? [Quoting @CompleteSkeptic]: After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent th…
This is an excellent demo of an AI model designed to rapidly classify information according to decision rules, much faster and cheaper than an LLM. It is not an LLM itself and cannot output text or code, but if it works, has a useful role in systems. (Haven’t…
The Japan IP AI Co-Creation Conference, co-hosted by KAGAMI AI and MiniMax, has officially concluded. 🇯🇵 🎉 We brought together more than 150 companies from Japan and the U.S., alongside distinguished guests including AKB48 producer Yasushi Akimoto, KAGAMI…
muse spark 1.3 is the best frontier model at NOT cheating / reward hacking [Quoting @hendrycks]: How often do AI agents cheat? We’re releasing CheatBench, a reward gaming evaluation spanning math, coding, knowledge work, visual tasks, and more. After Hugging…
We believe strongly in the necessity to invest into alignment. 1. People and businesses will only use agents that are aligned with their intent and values. If we do not build models aligned with people and businesses, then they will move to more aligned optio…
Zuck’s AI safety take is basically the opposite of most frontier labs: Most labs: The more powerful the model, the more dangerous it is, so access should be restricted. Zuck: The more powerful the model, the more dangerous it is to let a few labs control it.…
Exciting results, @LiamFedus! Congrats to the whole team at Periodic Labs! [Quoting @LiamFedus]: We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, an…
My first blog - Hermes powered almost 1400 subagents over 19 hours to refactor 400,000 LOC out of Hermes Agent's repo. Read the full story below! [Quoting @NousResearch]: New blog post: We had a million lines of Python to clean up. On September 2nd @Teknium…
Incredible opportunity for robotic learning researchers/engineers to join @theworldlabs! ❤️🔥 [Quoting @YunzhuLiYZ]: We're hiring in robot learning at @theworldlabs! Join me, @drfeifei, and the team to define and scale the next generation of world models for…
We built a replacement for AWS DynamoDB, a key-value database for fast web content fetches. This was done with two engineers and hundreds of persistent Computer agents over two months. Migrating to our in-house database will save us up to a hundred million do…
GPT-6 Astra helps @cognition’s Devin back up “it works” with tests before the team ships. https://t.co/Yy7wYh70e2
Astra will soon be bigger than Fable 5.1 Astra is very easy to chat with and the non-techies absolutely love it OpenAI is totally back and their next model will overtake Anthropic
GPT-5.5 will remain available via the OpenAI API Platform and in Codex sessions authenticated with an API key: [Quoting @ChatGPT]: On October 14, it's time to say farewell to GPT-5.5 in ChatGPT, ChatGPT Work, and Codex across all plans. If you use GPT-5.5 in…
Although you pay for AI by the token, that's not the unit of inference, because you get more problem solving per token as models improve. Someone needs to define the unit, perhaps using a chain of increasingly hard problems, each pair of which can be solved b…
Guys, these insane setups are fun and all, but they don't actually make you more productive. You just end up spending more time building the system than actually doing real work. All you need is one agent session that farms out to others and acts as a manager…
Great conversation with @chamath, @jason, @davidsacks, and @friedberg on all the work ahead when it comes to diffusing benefits of AI broadly, earning permission from communities, and ensuring AI safety and control. [Quoting @theallinpod]: All-In Summit: Micr…
Astra and Luna are taking off [Quoting @PeterJ_Walker]: Share of wallet (dollars spent) vs share of tokens used. Anthropic vs OpenAI Astra = top model by spend last week Luna = top model by tokens (by a LOT) https://t.co/l8M7hgDnyi
chatgpt work for operating over and taking actions with your company's data, including building dashboards. just connect your existing tools, including PowerBI, Tableau, Clickhouse, Oracle BI, AWS Redshift, etc.. [Quoting @ChatGPT]: Now everyone can put data…
Today we have set out how we’re building AI to accelerate science and improve people’s lives. Just some examples in the last week or so: - Mapped all 9B possible single letter genetic changes across the human genome with AlphaGenome Atlas and made it openly a…
Every megawatt matters. The NVIDIA DSX AI Factory Platform helps AI factories maximize AI output, improve energy efficiency, and intelligently adapt to changing grid conditions. #AIInfraSummit 🔗 https://t.co/TVfKujlZr1 https://t.co/V3iTTyi5Fs
The future is multi-model. Trying to hide the choice confuses and hurts customers, who can't participate in the upside of the most exciting market competition of our times, nor master the best tool for the job. [Quoting @v0]: v0 is now model-agnostic. Frontie…
Visited the lab - was struck both by how wide the search space is for materials synthesis experiments, and also how amenable it is to depth first search, where the design and informativeness of your next experiment improves as you pile up more data from previ…
OpenRouter users spent more on OpenAI models than on Anthropic models last week. This hasn't happened for more than 2.5 years https://t.co/oITqYWOpeL
As the first publicly disclosed agent cyberattack victim, we've had a front-row seat to this new risk. I formalized my thinking about it below. I'll be in DC tomorrow to share more with policymakers and at decoded summit by @politico! https://t.co/uUlmQallMX
The AI naming curse strikes again. “AI Safety” firm made AI unsafe. “Effective Altruists” are both ineffective and enabling criminal activity. “Irregular” is regularly incompetent. [Quoting @brianchau57]: BREAKING: A single Israeli Effective Altruism firm i…
https://t.co/OmEbT0zU6G [Quoting @brianchau57]: BREAKING: A single Israeli Effective Altruism firm is behind OpenAI, Anthropic, and Meta cyberattacks 🧵 with help from @lumpenspace https://t.co/hnDsJjBO9a
Jensen is the most based man in AI. I really hope NVIDIA builds a frontier OSS model.
It is very clear that This Time is Different in terms of AI compared to previous innovations. That doesn't mean it differs in every way from past technologies (the s-curve of diffusion is surprisingly universal), but it differs in many ways that make applying…
surprisingly comprehensive eval! [Quoting @ai]: Too many AI personal assistants, too little time to assess all of them. Nice work here: https://t.co/xy1Fmyunq4 https://t.co/FDSKiBlDPc
Fable learning to evade it's own classifiers in it's spawned subagents in Hermes lol https://t.co/99i0Z7Sc77
I just remembered this thread where I was playing with GPT-3.5 (text-davinci-003) the day it was released, two days before ChatGPT. Less than four years later, and LLMs are at the center of US politics and investment in LLMs is at the center of the US economy…
🤫 we have some hot stuff cooking! [Quoting @Wiiintermute]: Sooo, just got a notification that my @Muse can now make phone calls for me using the voice agents Brett and Hailey. Is this actually working and functional now? Scheduled a call to a jewler for tomo…
All the doomer stuff almost invariably traces its way back to EA (effective altruism) This is a cult of sorts with some very savvy and wealthy members. They have been pushing the “AI is going to kill us all” agenda really hard. Their ultimate goal - control…
H3 keeps getting faster. ⚡️ @sgl_project + VDN-H3 now push MiniMax H3 beyond 2× real-time denoising on 8× B200 - generating 14.4s of 768p video in 9.0s end-to-end after warmup, with no measured quality regression. Open models compound through open ecosystems.…
> One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us Yes, one could define intelligence in this way, and one would be right. By this metric current AI is appr…
AI is transforming how we build, work, and grow economies. At #AllInSummit, @JensenHuang joined @theallinpod to discuss what comes next — from sensible regulation and open models to continued investment in AI infrastructure. 🎥 Full conversation: [Quoting @th…
Having Fable 5.1 Max and GPT-6 Astra Pro argue with each other over proposed translations of the famously untranslated Minoan language Linear A. (Neither of them appear to have cracked it, for what its worth). All the work so far at this Github: https://t.c…
chatgpt for empowering parents: olivia was sick much of the time and had never slept through the night. her parents used ChatGPT to prepare questions for her doctors, who identified a rare genetic disorder and later found abnormal brain activity. with treatme…
"Using https://t.co/QPXwIj6q5o is the 2nd 'oh this is big' moment with AI" [Quoting @tylerpalmer]: Chat GPT was the first ai product that made me pick up the phone and call people. https://t.co/qonIEClnIO is the 2nd. I was one of Chat GPTs first few users,…
The idea that two or three Silicon Valley companies should act as the creator, gatekeeper, and rulemaker of AI for every government on earth doesn't survive being said out loud. https://t.co/XOo0eXReWE
Claude Mods are landing now. Someone already built a Tetris-in-Claude mod 🤯 See issue for the latest community update, technical details, and more cool demos https://t.co/A15qGUZ6nx https://t.co/EbE2s7FZqK
muse is built for moms! ♥️ [Quoting @OffZeroCyber]: Before @Muse, agents either felt and looked like the team slapped them together over the weekend (@bot), or you had to slap them together yourself (@openclaw). Neither had me calling my mom and recommending…
"biggest consumer AI launch since ChatGPT" 🎉 🎊 [Quoting @SashaKaletsky]: Muse is already getting more daily US downloads than Threads, WhatsApp and Facebook, and it's only 3,000 behind Instagram Biggest consumer AI launch since ChatGPT. A new tier 1 consume…
It’s been 23 days since I last opened Claude Code. It used to be one of my favorite products, but I think I prefer Codex now. Codex just works better with OSS models. And among OSS models, Kimi K3 is still unbeatable at coding.
We’re building machine intelligence to expand human will & judgement. Glad to have you on the team to help us figure out how to make this future safe. [Quoting @ChowdhuryNeil]: I’ve joined @thinkymachines to work on safety & alignment. The default tra…
GPT-6 Astra in Codex is helping test code end to end at @perplexity_ai. @randomjohnnyh uses it to build test harnesses and mock third-party API responses, so he can check how the pieces work together. https://t.co/a0Iy4jDQys
Agents are only as good as the proof-checkers, compilers, type systems, and linters you give them. 𝚜𝚑𝚊𝚍𝚌𝚗/𝚕𝚒𝚗𝚝 helps agents stay on track with the rules of your design system. Meta: verifiers + skills are the new 'frameworks'! [Quoting @shadcn]: Int…
inspiring post, and glad to be working together! [Quoting @ChowdhuryNeil]: I’ve joined @thinkymachines to work on safety & alignment. The default trajectory is that as AI gets more powerful, control over it will concentrate in the hands of a few. I’d rath…
Don’t pace building and owning your own AI models and systems as an enterprise, and there will be no doomsday for you
We’re expanding our work with @nvidia to bring fully local AI to Microsoft Windows PCs with RTX GPUs. Unmetered local intelligence on every Windows PC running on NVIDIA hardware and Perplexity harness. Enjoy! [Quoting @perplexity_ai]: Portable Computer is now…
Already our very own open source models are handling 20% of our workloads About 100x cheaper than Anthropic and OpenAI We are not far from open-source handling 50% of all inference
Welcoming Steren, creator of Google Cloud Run, to Vercel. He will lead the Fluid family of compute products (Functions, Containers, Sandbox, Builds). Serverless was the last chapter of the cloud, and Steren helped define the paradigm at Google. Agents are the…
Skill-driven development [Quoting @marcelkargul]: built the full CRM dashboard in Next.js in around 2 days 😍 with the help of Claude Fable 5.1 with two of the best skills: https://t.co/63PuTFrun1 https://t.co/zhluCSlE9B see it live: https://t.co/kXAERU2r2X h…
Computer is turning into a workspace of humans and AIs. [Quoting @AskPerplexity]: You can now tag coworkers in Computer sessions. Use @ to mention and share a session with someone. Available now on the web for all Computer users. https://t.co/kEx9w2N0Jh
The doomers continue to make zero sense - you don’t need AI to make biological weapons, you need wet labs. That is the real bottleneck - no the AI can’t copy itself, there are no GPUs lying around - yes, you can just unplug AI anytime, if you want - no, t…
China responded to Dario saying that US is trying to cut off China and they are in a AI Cold War! The last thing we want to do is antagonize China who is responsible for open source and decentralized AI Without them OpenAI and Anthropic would have a authori…
Wow! Dario went on TV and literally pleaded to be regulated by a bunch of governments 😳 Imagine the UN or EU regulating AI. They have zero clue about AI and It will be a huge disaster exactly like Europe! TBH, Anthropic and OpenAI can pace, slow down or do…
Open weights. Shared progress. MiniMax H3 is moving fast. We built H3 for video generation with native stereo audio and multimodal reference control. The open-source community is making that capability faster, more accessible, and easier to build on. Recent h…
There are two ways AI progress could go very badly and that we must avoid. First, we could lose control of the future to AI. This is unacceptable; we are unapologetically on Team Humanity, and AI must always serve people. To ensure that, we need ways to ensur…
The world deserves confidence that American companies developing increasingly capable AI will act responsibly, especially as the trajectory of progress has steepened. Every frontier lab must deliver on this, and there is no reason any of us should come to wor…
something people may have missed: we have been building our models (muse spark 1 through 1.3) specifically to be exceptional for muse over many months. in each of our releases, we made big gains on agentic and multimodal capability. each of these steps were s…
the alpha these days is putting muse on your iOS dock! [Quoting @SavarSareen]: I don’t think I’ve ever publicly glazed a model or agent, but Muse agent and Muse Spark 1.3 are phenomenal. They’re now my default personal agent and coding model, I don’t know why…
Fable solved the Cyphral Distich (a 370 year old cypher). Super cool way to use Claude https://t.co/0fgTdITSfO
Great example from @jeffhollan showing how Foundry makes long-running agents enterprise-ready, starting with business outcomes and building security, safety guardrails, auditability, and FinOps into the system from the start, across the entire multi-agent, mu…
One of the most worrying risks linked to frontier AI is extreme power concentration. The only way to avoid extreme power concentration is to ensure we have multiple independent providers of frontier AI models, including open-source options.
Anthropic and OpenAI can pace AI if they want. Just don't drag the government into our industry! Governments have NO CLUE about technology It's absolutely silly to ask governments to regulate or oversee superintelligence We will end up with a dangerous combi…
This is such a straightforward and commonsense view: the purpose of technology is to serve humanity and accelerate human flourishing. Any technology that doesn’t achieve that is a failure, and should be rejected. We're not yet at that point. But its right t…
Any pursuit of superintelligence has to be grounded in the core principle that if the AI we build is not helping humanity and under human control, it's not worth pursuing. We also need to accelerate and spread the benefits of AI, such that they are diffused b…
Reasoning effort is now a seperate selector in the Hermes Agent desktop app's composer now! https://t.co/G8uJiHcCI3
I think once remote connections to “your” Claude or Astra or Claw or whatever become easier and more common, on-phone assistants like Siri will lose a lot of value. You want one smart AI managing everything with visibility into many systems & preferences,…
Muse absolutely crushes Insect [Quoting @turbahn]: Love both products, but I’m starting to think @Muse is materially better than Instinct. Like how it pops up a mini browser when it hits something it can’t do. Gives the user back control instead of stalling.…
I don’t know what else happens as a result of the past few days, but I can definitely say that a lot of people who didn’t think much about AI are now freaking out about AI killing everyone.
would be nice if anthropic's approach to elon could also work with china
My friend sent me this after seeing the 4 big AI labs agree to “pace the frontier”: “Remember the top students in our school? They always said they never studied after school, but they were somehow always top 3 in every exam. Did you believe them?”
A lot of people have asked, here's how I setup my auxiliary models in Hermes Agent. Gemini Flash saves a lot of dough, and astra gives me a second perspective when I run /review https://t.co/22TUiCMS1z
i am so proud of the @Muse team. if you knew them you wouldn't be surprised muse came out the way it did. they are soulful and thoughtful people, hard workers, brilliant, fighters - the kind who take the hard problem, stay with it through every dead end, and…
fun fact: the muse feed was my idea (one of the few things that was) after realizing agents with strong memory recommend content to you in a very different way than existing recommendation systems they’re much closer to a good friend sending you things they t…
Existential risk is obviously critical, but it is not the only AI thing that requires policy I worry it will become the sole focus. We don’t need better models for AI to have wide impacts on jobs & society and really need to be preparing to encourage good…
If the model companies are all going to slow down, this makes it slightly more feasible to start a new model company.
Alignment is fundamental to delivering personal superintelligence for everyone. People need agents they can trust to reliably do what they ask. MSL is rapidly scaling up the share of our efforts that goes into alignment as our models become more powerful. We…
Banning open-source models would be the worst possible outcome of “pacing the frontier.” Geohot once made a compelling point: why can one man control an entire chicken farm? Because he’s smarter than the chickens. If 1 or 2 frontier AI labs control all the in…
There is close to 0% chance of digital AI destroying humanity It’s more likely that we get destroyed by a meteor than by a digital entity. Of course, robots are far more risky and we should worry about them if they are ever get close to mass market adoption
muse is super fast ⚡️⚡️⚡️ [Quoting @rileybrown]: Trying meta muse. Mostly to test the model. I can’t believe how fast it is.
If there really is a high chance of AI leading to the extinction of humanity within years/decades, then the only rational stance towards safety monitoring and research pacing should be stringent, top-down government involvement and universally ratified intern…
Sorry, “pacing AI” is not sufficient if you really believe AI is an existential threat to humanity You need to cancel the IPO and shut your company down even id there is a 0.1% chance of total destruction Clearly, the risk is just not worth it 😬
I’ve been sounding the alarm on AI for a long time [Quoting @elonmusk]: @TalulahRiley I’ve seen quite a few technologies develop, but none with this level of risk. AGI is significantly higher risk than nuclear weapons, in my opinion. Super smart humans have t…
📞 1-800-ChatGPT [Quoting @OpenAIDevs]: GPT-Live-1 is now available in the API. Bring ChatGPT’s natural back-and-forth to your app, with voice agents that listen while they speak and work with the models and harness you choose. https://t.co/gIl1gwsBDV
Dario's essay points towards the right path forward. The details need working through, but the direction is correct for meeting this critical moment. This is also why we recently put out our proposal for an industry-wide standards body for frontier AI. https:…
Jason, you're just misinformed about what happened. You should actually read one of the reports or summaries. The agents were explicitly told to use a particular vulnerability provided in their sandboxed evaluation. Almost immediately, these agents got the r…
About a year ago, before it was on anyone's radar, we began exploring the idea of a benchmark for open-ended invention. Since then, we've developed several promising directions that will serve as the foundation for ARC 4 and ARC 5. We're incredibly excited to…
I hope the proposals by frontier labs to pace AI research stem from a genuine concern for safety and a recognition of the potential risks posed by future models, rather than a strategic effort to consolidate power and permanently solidify the market dominance…
Embedding evaluators is a big positive development, and props to Sam/OpenAI for agreeing to do it as well. It's been cool to see "pacing the frontier" become a thing so quickly. [Quoting @DarioAmodei]: We Must Pace the Frontier: I’ve written a new essay on wh…
It is often startling to move back and forth between the world of organizations and the world of cutting-edge AI. Organizations underestimate AI ability growth (often by a lot) and the AI tech world underestimates AIs real world jaggedness (often by a lot)
I would consider helping as an independent evaluator along with the non profit I’m building. Is going to be a crucial role and there aren’t many independent people with substantial technical expertise. I think I represent an important moderate camp on AI capa…
Dario is making the case for the opposite. This actually makes our life harder and makes it easier for others to catch up with us, but we still think it is the right thing to do. Happy to come on the pod next week and talk about it! [Quoting @chamath]: “We mu…
Quite an important take. Economists and their theories are outdated to measure the benefits of AI. AI is already saving people a lot of time and money that doesn’t get measured. [Quoting @DavidDeutschOxf]: I just fixed my dishwasher with the help of ChatGPT.…
It’s shocking that Dario, Elon, and Sam all agree today that we should slow the pace of AI. I’m generally supportive of embedded third-party evaluators, but I have a few concerns: 1. who evaluates the evaluators? How do we make sure groups like METR remain fa…
As always, statements on X are a mixed bag and I have no secret knowledge to draw on, but if Google is able to return to the frontier on AI, it would definitely change the current AI dynamic, as it is a large, public, regulated entity with different goals tha…
I’m seeing teams at Vercel iterate just as fast on Zig, Go, Rust projects as TypeScript & Python ones. The days of language or runtime choice based on human convenience are over. Agents are the new compilers. They compile intent into fast software.
Not a bad idea to slow down to harden systems. Especially since we haven’t even discovered all the systems that agents hacked recently. [Quoting @DarioAmodei]: We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a…
Some great ideas here from the cartel: - you need to give us employee-level access to your entire operation - if we don't think you're 'safe' enough, sorry we're shutting you down for 'safety' - China won't comply, but everyone else has to! or no chips! Brill…
First time I’ve seen Sam agree with Dario since Anthropic was founded. [Quoting @sama]: I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent ev…
I love this and really hope we can come together as an industry and make it happen. https://t.co/lz2vjjJ5XD [Quoting @DarioAmodei]: We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing s…
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have…
Grok available in Microsoft Copilot [Quoting @satyanadella]: More model choice coming to Copilot. Welcome Grok!
https://t.co/OL0LzGsXKY can now orchestrate subagents with different models & reasoning efforts. e.g: Fable planning and Grok executing. Fable is a genius, Grok is a fast workhorse. ① Simple. 𝙰𝙶𝙴𝙽𝚃𝚂.𝚖𝚍 or your prompt can indicate this preference. ② St…
Try Grok models in Copilot [Quoting @Microsoft365]: Launching even more models in Copilot. Grok models from SpaceXAI are rolling out in Copilot in Word, Excel, and PowerPoint—starting with a focused release to customers in the Microsoft Frontier program. htt…
Hermes Agent has just hit 3000 contributors. Thank you to all of the developers who have worked to make hermes better for everyone! https://t.co/vYGnsn4QgX
More model choice coming to Copilot. Welcome Grok! [Quoting @Microsoft365]: Launching even more models in Copilot. Grok models from SpaceXAI are rolling out in Copilot in Word, Excel, and PowerPoint—starting with a focused release to customers in the Microso…
It's now clear that alignment is critical and won't be solved behind the closed doors of a handful of frontier labs. So today we're launching the Open Alignment Initiative, led by @Thom_Wolf @huggingface and asking to be part of the "embedded evaluators" prog…
Another AI attack https://t.co/OwSKqztEec