What happened in AI and what people make of it. Every number measured at the primary source.
Anthropic just killed the Cowork name two weeks after OpenAI killed the Codex name, but nobody is asking why both labs are suddenly ashamed of the word "work."
On September 16, Anthropic folded Claude Cowork into plain Claude. The blog post is a single paragraph, mostly about Simon Willison being relieved he does not have to write a guide explaining the difference anymore.
Here is the buried part: two weeks earlier, OpenAI renamed its Codex desktop app to ChatGPT. Two different labs, same move, same month. The Codex rename got a full blog post with screenshots. The Cowork merge got a link-blog shrug. But the pattern is identical — both companies are backing away from naming a tool after the thing you do with it.
@Sherveen spotted what is actually being lost: if you ask a thinking question in Chat versus Work mode, you get a different answer. Merging them is not just a UI cleanup. It is a choice about what kind of thinking the model does by default, and you no longer get to pick.
For anyone using these tools daily, this means the "just chat with it" mode is eating the "actually work with it" mode. The agent that edits files, runs commands, and ships things is being hidden inside the same box as the one that tells you fun facts. You will have to guess which one you are talking to.
My bet: this is not about simplicity. Both labs are racing toward the same interface — one box, no modes — because the next step is a model that decides for itself whether to chat or act. They are not removing the work mode. They are removing your ability to know when it is on.
the receipt “If you ask a thinking/research type question in 'Chat' versus 'Work' mode in these products -- say, something complex about politics, or…” — @Sherveen
someone put it well@Sherveen: People keep asking for this w/ Codex, too, and I really regret that both labs seem inclined to listen. If you ask a thinking/research type question in 'Chat' versus 'Work' mode in these products -- say, something complex about politics, or…
the commentsThe core tension is between excitement for a vendor-neutral GPU programming future in Rust and deep distrust of NVIDIA's proprietary lock-in history.
I strongly dislike CUDA. Once you have allowed that proprietary cr*p into your C++ codebase, it is very hard to get rid, and you end up with code that is either tied to a single vendor or an #ifdef hell, probably both. The best way to prog…
Since NVIDIA owns huggingface now and huggingface has the excellent Candle [1] crate for inference on Rust, this seems like a good step towards nice native Rust kernels. [1] https://github.com/huggingface/candle
Really exciting but it reads like Claude instead of what Nvidia posts have generally been like in the past. I don't need nor want my tech blogs to sound like a young adult novel.
Anyone know when Rust's std::autodiff will become stable? Assuming this Rust support expands to other GPU vendors, autograd will probably be the only reason to use Slang instead of Rust anymore.
the post saysTypeSafe AI announces Jev, a model using Reinforcement Learning for Calibrated Decisions (RLCD) that guarantees 0% type errors.
the commentsThe core tension is whether this model's speed and cost advantage for structured queries justifies trading away the open-ended generation capabilities of a general-purpose LLM.
First, congrats to the team on launching something genuinely interesting and new. Seems like a more accurate title would be "Jev: Trading general purpose generation for fast typed inference" or something like that. This is interesting, but…
Wasn't really till seeing this home assistant demo they have (https://www.loom.com/share/18c4dbcf8db546dfb2d7f2ef018e78e4) that the value really clicked for me. Seems really cool.
This, combined with contracts, could make a lot of things so much fun now! For those who don't know (which is probably everyone but me), I ported the design-by-contract pattern in Python and combined it with LLMs. This was early 2025. I or…
This is a very promising idea - a model that takes arbitrary text input (which can be a complex json), plus a set of questions (yes/no, multiple-choice, or score) and quickly (milliseconds) and cheaply ($0.042/MTok) answers those questions…
“OpenAI confirms weeks of AI safety talks with Anthropic and Google DeepMind, as Trump's team dismisses safety concerns and pushes to keep pace with China.”
@bikepedantic.bsky.social
if this is just an excuse to curtail their runaway cash-burning, this is just straight-up collusion techcrunch.com/2026/09/15/o...
the commentsThe core dispute is whether the claimed speedup is a genuine advance or a brittle result from overfitting to an unrealistically small, in-memory dataset.
“81% faster query plans than Postgres”…on an 8 GB dataset that fits entirely in memory, with shared_buffers constrained to a fraction of that, queries warmed before measuring, and read-only SELECTs. I would be cautious about over fitting, …
Engineer: "HELP, our production DB is frozen on this query that worked fine before!" Infra: "Hmm, let's check... Well would you look at that, it seems like your LLM query planner usually works and produces fast queries, but this time when …
Optimal plan construction is math-heavy, algorithm-heavy and vary even by workload. There are options like creating just-in-time indexes, so solution space grows even faster than article presents. Sometimes it is the query planner which is…
> Frontier intelligence is extremely powerful; the distillation I did off Astra trajectories is proof enough that large models are not going anywhere Wouldn't admitting this invite trouble due to accusations of distillation flying around b…
the commentsThe core dispute is whether inserting ads proves OpenAI is desperate for revenue rather than confident in its product's intelligence.
“Explore new AI-powered advertising experiences from OpenAI, including Sponsored Agents, tools for marketers, and integrations with HubSpot and Shopify.”
“OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.”
@timkellogg.me
openai wrote up 6 misalignment issues with lots of examples openai.com/index/model-...
the commentsThe core debate is whether Xiaomi's rapid open-source progress threatens proprietary AI companies or simply reflects incremental engineering gains on benchmarks.
I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothin…
Neat! I've been trying out their next model for the last week, which I assume is a version of this, and it's been a good experience so far. I had used 2.5-pro for a hefty chunk of development, and found it to work like a somewhat forgetful…
For reference, Mimo-v2.5-Pro scored 19% on DeepSWE 1.1. This is looking great. Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort). https://deepswe.datacurve.ai/blog/deepswe-v1-1
the post saysAnthropic trained Claude on a constitution stating its moral status and consciousness remain deeply uncertain.
the commentsThe core dispute is whether an AI's convincing performance of self-preservation indicates genuine sentience requiring moral consideration or is merely a sophisticated mimicry of human training data.
I appreciate his openness. > Unfortunately, there’s a growing chorus of people who argue that AIs could now be, or may soon become, conscious. They argue that AIs may deserve rights and protections similar to those that we provide other co…
Birch, The Edge of Sentience (2024), ch. 16 - "simply no way to assess sentience in an LLM" Schwitzgebel, AI and Consciousness (2025) - "we won't know before we've already manufactured thousands or millions of disputably conscious AI". But…
Look, I do not have a scooby if current AI models are conscious and I strongly suspect it’s a meaningless question, but sooner or later we will need to address whether or not a certain thing is or isn’t a person, and we’d better not screw …
The problem with this premise is that models are trained on a vast corpus of human behavior, which they emulate with varying degrees of effectiveness. Humans, unsurprisingly, act as if they have a stake in their own well being, value their…
the post saysMistral powers Firefox Smart Window in France and North America, with the UK and Germany planned for later in 2026.
the commentsThe core dispute is whether browser-based AI should run locally for privacy or rely on cloud processing despite claims of privacy safeguards.
This is an excellent use case for completely local, small model inference, yet for inexplicable reasons Mozilla wants to normalize uploading your entire private browsing history to a cloud. These (Mistral's and Mozilla's) marketing pages a…
Would be very cool if you could type a long query and the model would just build an advanced google search query using what you typed. You could even ship a tiny model in the browser itself that does that. Something like: "blog posts which…
It's interesting to see Firefox attempt to create slightly more privacy-focused cloud inference infra (assuming you can trust that they adhere to their own policies and don't have bugs, and that their partners adhere to their contractual o…
Seems like what Chrome has with the default built-in Gemini Nano model. Hey at least "some" news/things from Mistral. Seems like ages ago when they launched vibe-code. > Powers context-aware search, page summaries, and memory retrieval acr…
the post saysThe last version of PS5 Linux helmed by the developer is Version 2.5 supporting PS5 Phat and Slim consoles on firmware 3.00-7.61.
the commentsThe core dispute is whether the project collapsed due to an influx of low-effort LLM spam drowning out meaningful collaboration, or because a key exploit was leaked in violation of an embargo.
Hobby groups projects like this are less fun for a lot of people who used to enjoy interacting with smart people. It's definitely become a game of just spam claude for answers with zero understanding or care for how anything actually works…
now imagine how many other devs and maintainers are contemplating this but just haven't been pushed to their personal breaking point yet. or take note of how many, when asked about the spam problem, just nervously go "yeah it's kinda rough…
Person 1 finds an exploit the old fashioned way. Keeps it secret because they’re doing some PS5 Linux work and want to keep it unpatched until GTA6 releases in 2 weeks. I don’t have a PS5 but I assume new games might require a firmware upd…
Misleading title, there's another major reason: it's an embargo agreement violation that jeopardizes Linux support on PS5. https://x.com/theflow0/status/2099987019954831744
“Last day to book your exhibit table at Disrupt is September 18. Three days left. Get your startup in front of 10,000+ founders, investors, operators, and tech leaders on October 13–15.”
This is the most grounded and coherent take I’ve seen on the actual realizable value of LLMs.. pretty much since they came out. > the classes of firms that can accept the use of fully autonomous LLMs are few, by my count just three: 1. tho…
This April 2026 paper is a fun and related read. https://arxiv.org/html/2509.24239v4 Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identi…
The premise in the very first point seems off: > the frontier labs are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers... Ev…
Great article, but the lack of sentence capitalization makes it unnecessarily difficult to read. Apologies if this comment is off-topic, but it really is quite egregious, and since the article was submitted by the author I presume they are…
I too wrote about this, although _much_ longer, if you want a bit of a deeper dive on the subject https://www.iankduncan.com/personal/2026-09-16-sex-ai-and-th...
I've tried to read this, but somehow the article doesn't really come to a point. Is there any point besides Yud not being perfect, and having connections to people who openly talk about their sex life (and that it's somewhat non-standard)?
Any actually-reasonable person who was around for the founding craze of this era, ie., New Atheism, quickly realised that there was a deep pathology at the heart of many of its zealots, perhaps summed up by a desired to "recreate religion …
the post saysGemini 3.8 Live Extended Thinking achieved the top score of 82.6 on a speech-to-speech quality evaluation.
the commentsWhether Gemini is actually reliable and competitive or still lags behind despite Google's resources.
@five-oclick-tech.bsky.social
Gemini 3.8 Live for smarter voice agents 🤖 Google unveiled Gemini 3.8 Live models, boosting near real-time reasoning for voice agents that can now think and talk simultaneously. Image: DeepMind Blog
the post saysClaude merges Cowork and chat so that a single conversation can produce a report as a doc and matching slides, both shareable via one link.
the commentsThe core debate is whether merging chat and work modes into one interface is a genuine simplification that empowers users or a misguided design choice that degrades focused research and encourages unrealistic, always-on productivity fantasies.
People keep asking for this w/ Codex, too, and I really regret that both labs seem inclined to listen. If you ask a thinking/research type question in 'Chat' versus 'Work' mode in these products -- say, something complex about politics, or…
How did this new account `vertigoruntime` get three posts on the front page, all in the last day? - This post - The DeepMind Institute https://news.ycombinator.com/item?id=49727659 - Mistral X Mozilla: Private, Multilingual AI Browsing htt…
These kinds of updates always have this romantic scenario of someone having Claude develop a presentation or something on their way to work between multiple devices, which actually feels a little sad and does not align with what happens in…
Hi, this is my team! Happy to answer any questions. There's a lot in this launch, but the core idea is to simplify the product while giving users access to more capabilities. You no longer need to know ahead of time how much work a convers…
the post saysGoogle launched Gemini 3.8 Live Extended Thinking, which scored 82.6 on the Speech to Speech Quality Index and 68.6% on the τ-Voice benchmark.
My first language is Afrikaans, which is a somewhat niche language and hard to find teachers/conversation buddies outside South Africa. (I live in USA now) I've been using Gemini to live chat in Afrikaans and do impromptu Afrikaans grammar…
Just gave it a try - very solid release. Copes well with thick accent, voices are pleasant and latency seems low. Oh and I can actually use it on a workspace account - which for most of the recent releases was an account stuck in limbo. No…
I don’t understand good experiences people are having with Gemini. It’s the only model that sometimes loses/forgets context in literally next message. Plus feeding unasked product links to responses.
I wonder when/if we’ll see Gemini beating Fable and Astra. Last year I would have confidently bet Google will overtake the others just because they have the data, the hardware (TPUs) and a fat advertising money pipe and yet they are still …
the post saysIBM introduces the Consistency Analyzer, a diagnostic that resamples a single recorded trajectory to identify flip-prone decision points using one additional model call per step.
Interesting to see these experiments. This is early but imagine in few years there will be companies mostly run by agents with a light overview from a human operator. What then happens to scaling of the bussinesses? I would assume, just li…
Meanwhile, I can't even get Astra to consistently re-use the same font-size across all of my HTML page headings+subheadings. There's just no world where this actually results in a stable, respected business. It will be death by a thousand …
I am in the process of attempting to have AI run my business. I'm actually making very good progress, but it's happening in pieces - I document some task and have it take over, or I give it something to handle while staying in the loop and…
There was surprisingly little information on how they actually do this, but we run our business with a large number of, what we call "AI employees" in addition to regular employees, and they act in interesting ways. We've been building out…
“The true measure of AI is who it helps. Here’s how it’s impacting lives today. We're focused on key areas where advanced technology can help make extraordinary progress …”
Dear fellow humans from "Hacker News". Hacking a driver that in itself documentation to black box Apple hardware is not any different from hacking $10 4G LTE modem. Fact that a person who was not previously driver developer can achieve thi…
https://www.reddit.com/r/AsahiLinux/comments/1whecn1/comment... > The author was banned from Asahi Linux for hiding his extensive use of LLMs from us in another attempted contribution, and (more importantly) for concealing that he is a for…
It's extremely impressive that they were able to make a working driver so quickly. I think this is one of the best use cases for LLMs. You don't need someone to spend years reverse engineering undocumented hardware anymore. It will interes…
This is super great. The biggest pain point of Asahi Linux is how it doesn't have GPU acceleration on M3 and newer, especially now that M6 is out! However, Asahi Linux has a strictly no-AI policy [1]. So this great work can't be upstreamed…
The Irregular post mortem comes down to lack of basic security controls "Ultimately, most of the issues we’ve discovered were due to internet access controls." That seems so incredibly basic and common sense that you would test and monitor…
My understanding is that Irregular were the company that hosted sandboxes to run some of these evals in, and those sandboxes ended up misconfigured. I got the impression that in some cases it was the customer (Anthropic etc) misconfiguring…
Nevo was in Unit 8200 for years. Companies started by Unit 8200 members always have mysterious exploits like the vibe coding Wix exploit. So either it was a deliberate exfiltration channel for e.g. getting the entire model or they were in …
Anyone else walk away from reading this, and looking at other articles there, and get a weird feeling about that site? Like, there was some weird stuff about migrants and gender equality, and stuff about Islamic terror; mixed in with some …