Microsoft's GitHub source shack is reporting a 50 percent error rate on repository content downloads. GitHub Copilot is also "experiencing degraded availability." The problems began at 1340 UTC, when the company's status page reported "impacted performance for some GitHub services." This rapidly escalated to degraded performance across numerous services, including Issues – which was, appropriately enough, experiencing an issue of its own. GitHub is investigating and has promised an update, but today's incident is only the latest in a string of disruptions that have left developer nerves frayed. Earlier in August, GitHub's Actions automation platform and Pages hosting service experienced problems. In its May 2026 availability report, the company acknowledged that AI-assisted coding and agentic workflows were adding to the strain. GitHub parent Microsoft boasted last year that AI was writing as much as 30 percent of the code in some of its repositories, subject to human review, according to CEO Satya Nadella. Addressing GitHub's repeated outages in June, software engineering SVP Jakub Oleksy said structural changes were afoot to "permanently remove failure modes." "We acknowledge that we have work to do, but we're committed to getting it done and making GitHub reliable when and where you need it." Brave words, but today's disruption demonstrates that there is still plenty of work to do – and each recurrence further tests developer confidence in the service. Unsurprisingly, developers took to social media to vent their frustration as Monday Mondayed even harder thanks to the issues. One posted, "Github down again, abou[t] as surprising as the sun rising." Another was more plaintive: "Github is really falling apart. Can we stop being down please... We got work to do." The cause remains unknown, but GitHub's disruptions are stacking up, leaving developers to count the idle hours and consider alternatives. "Maybe," pondered one developer, "the AI code submissions are too massive for them servers." Maybe they are. ®
The cost of agentic AI workflows is forecast to increase more than fivefold by the end of 2028 as users adopt more complex applications of the technology. As Nvidia and other tech giants push inference and agentic AI as the next stage of the AI wave, Gartner warns that the cost of implementing these systems will rise even as foundation models become cheaper. The analyst firm has cast its eye over the nascent world of AI agents – systems designed to act independently in pursuit of a goal – and sees multiple challenges ahead. Leaving aside the substantial security concerns, these software agents are considerably more complex than chatbots. Gartner believes falling model prices are tempting users to build more complex workflows, whose greater token consumption can outweigh those savings and drive up overall inference costs. In other words, tokens are becoming more cost-efficient, but those savings are not keeping pace with the rising cost of more advanced AI capabilities. The rate of innovation is outpacing the cost curve, Gartner claims. "The harsh economics of the inference paradox are exemplified by the differences between a simple chatbot and an AI agent," says Gartner senior director analyst Will Sommer. "Where a simple chatbot must read and interpret a query and quickly respond with a probabilistic reasonable answer, an AI agent must constantly reason, negotiate, and question itself," he explains. Those processes add up: routing a task to an agentic reasoning model increases inference costs at least fivefold, and potentially by much more as the task becomes more complex. Securing a return on investment from such advanced AI tools therefore demands either much greater returns than basic models provide or better optimization of inference, routing, and orchestration, Gartner warns. This could mean assigning each task to the most cost-efficient model capable of handling it. The move by some AI providers from flat-rate subscriptions to usage-based billing hasn't helped, as The Register reported last month. Token-heavy workflows can produce runaway costs under the new model. Perhaps it is no wonder Gartner predicted earlier this year that 40 percent of organizations would demote or decommission AI agents because of problems with the heavily hyped technology. The analyst biz also cheerily forecast last year that at least half of all generative AI projects would blow their budgets because of poor architectural choices and a lack of expertise, while most attempts to build custom models would be abandoned. ®
Microsoft will retire Excel's COPILOT() function on September 14, barely a year after its preview debut. Introduced in August 2025, initially for Beta Channel users with a Microsoft 365 Copilot license and later for Excel for the web users through the Frontier program, the COPILOT() function let users send instructions to the company's AI assistant directly from a worksheet cell. Microsoft has now decided the Copilot side pane should be enough for anyone. Beginning September 14, 2026, the COPILOT() function will no longer be available. Microsoft had planned to make the function generally available in 2027. It has now updated the Microsoft 365 roadmap to say: "We have decided not to move forward with this feature. We apologize for the inconvenience." According to a Microsoft 365 Message Center post, "this change helps streamline the Copilot experience within Excel while continuing to provide AI assistance through a supported interface." "Customers can continue using Copilot in Excel through the Copilot side pane, which provides many of the same AI-powered capabilities, including summarizing text, classifying data, generating content, and retrieving information from the web." Since the function remained in preview, spreadsheet wranglers should not have relied on it in production workflows. Anyone who did now has some formulas to replace. In its original announcement, Microsoft wrote: "Its output should be reviewed and validated for accuracy, especially for critical business decisions or reports." Hardly music to the ears of spreadsheet users, for whom accuracy tends to matter, but par for the course with AI services. Microsoft has spent recent months trying to make Copilot behave more consistently across its productivity applications. It has also responded to some user feedback – the Dynamic Action Button can, for example, be banished to the toolbar rather than left floating atop the content. With similar capabilities remaining in the side pane, Microsoft evidently decided that a dedicated worksheet function was surplus to requirements. Google Sheets still offers a similar AI function. The search giant announced the feature in June 2025. ®
UK supermarket giant Sainsbury's temporarily suspended live facial recognition (LFR) at one of its stores after staff wrongly ejected a shopper in response to an alert – the second such incident this year. Matt Arnold, 46, was ejected from the chain's store in East Dulwich, London, on August 6 after staff responding to a Facewatch LFR alert apparently mistook him for a suspected thief. The comedy promoter told the BBC that he was using a self-service checkout and called an assistant over to approve an alcohol purchase. Instead, two Sainsbury's managers approached him, refused to serve him, and linked him to an incident earlier that week. Arnold said he was told that the store's LFR system had identified him in connection with a previous offense before he was escorted from the premises. "They came over and said I had to leave," Arnold told the broadcaster. "The staff member said I'd been identified by the AI, and the cameras had flagged me. "A shoplifter does not walk around with that much shopping, they don't scan it through, they don't put their Nectar card through. "But what upset me was thinking this is what the future could be – people just listen to what the machine tells them to do without thinking about the consequences." Sainsbury's told The Register that the mistaken ejection resulted from human error rather than a false match by the Facewatch system. A spokesperson said: "We have contacted Mr Arnold to apologise for his experience at our Dulwich superstore. "The incident was caused by human error, not the facial recognition technology. Customers can be reassured that the Facewatch system has a 99.98 percent accuracy rate, and every match is reviewed by a trained manager." A Facewatch spokesperson likewise said: "We can confirm that our live facial recognition technology was not at fault in this incident. A correct alert was sent to the retailer, but was subsequently subject to human error in the way it was handled and communicated by the retailer." We asked Sainsbury's what exactly the human error was in this case, but it did not reply. The Register understands, however, that Arnold's ejection was similar to that of Warren Rajah, who was wrongly removed from the chain's Elephant and Castle store earlier this year. As in Arnold's case, Sainsbury's maintained that the Facewatch system had worked as intended. Regarding Rajah, the supermarket said the system correctly identified someone linked to a previous theft, but staff responding to the alert approached the wrong person. Rajah said at the time that he was approached by three store managers holding smartphones. They looked at the screen, then at him, told him to leave, and pointed to an LFR flyer posted up near the store's entrance. Arnold told the BBC that, as he left the store, he looked behind him and saw what he believed was a security alert displaying his face inside a red circle. Sainsbury's temporarily suspended LFR alerts at the store while it reviewed its internal processes and considered additional staff training. Facewatch said: "When we suspend a retailer from our system, this is a precautionary measure that prevents the store from receiving further alerts while the retailer reviews its use of the system and undertakes appropriate actions, including further colleague training where necessary." Sainsbury's announced a major expansion of its LFR deployment last month, confirming plans to install the technology in up to 200 stores by the end of the year in an effort to tackle shoplifting. The technology currently operates in 55 stores, and the supermarket claims that 90 percent of people identified through the system do not return. Despite Facewatch's oft-touted 99.98 percent accuracy figure, other wrongful interventions involving the technology have been reported at UK retailers. Retailers using the technology include B&M, Budgens, Costcutter, Iceland, Southern Co-op, Spar, and Sports Direct. ®
In an effort to "watermark" text that Claude has generated and comply with the EU AI Act, Anthropic unveiled a plan on Friday to modify its bots' choice of words in a way that would be detectable as the product of an AI. Traditional watermarks are patterns or images overlaid on currency, postage, or official documents as an assertion of authenticity. In the digital realm, the term is more flexible and can refer to a variety of techniques for applying an identifier to electronic data. Anthropic's approach involves influencing inconsequential word choices made by its models, a technique introduced in Google DeepMind's SynthID-Text paper. To oversimplify things, large language models are fancy autocomplete engines which work by predicting the next word in a sequence of words. Anthropic explains that while composing sentence output like "The weather today was cold and…" a model like Claude might respond with words like "cold" or "gray" and would be unlikely to respond with a word like "sugary." That's the theory, but when actually asked to complete that sentence, Claude Opus 4.8 went a bit overboard: "…crisp, the kind of cold that nips at your fingertips and turns your breath to little clouds. The sky was a pale, washed-out blue, and everything felt sharp and clear." And then it checked to see if users thought that was useful, asking, "Want me to take it somewhere specific — cozy, gloomy, cheerful? Or keep going with the same tone?" But remove whatever training has been applied to promote engagement and simulate literary style, and that's basically what Claude is doing here – predicting the next word in a sequence. Anthropic asserts that in most cases, the example sentence could be completed by either "cold" or "gray" and "the meaning of the sentence is largely the same either way." The watermark is generated by deviating from the predicted word to something else. A different source of randomness is used and that can be detected with a digital key. As Google DeepMind researchers explain in their paper: "Generative watermarking works by carefully modifying the next-token sampling procedure to inject subtle, context-specific modifications into the generated text distribution. Such modifications introduce a statistical signature into the generated text; during the watermark detection phase, the signature can be measured to determine whether the text was indeed generated by the watermarked LLM." Anthropic insists this will be done with low-stakes passages in a way that won't alter the meaning. "In internal testing, we’ve seen no impact of watermarking on the content, level of creativity, or readability of Claude’s text," the company said, adding that in a controlled study, human raters saw no difference in quality between watermarked and unwatermarked answers. That assumption hinges on not applying the watermark to any consequential text. As Anthropic puts it, "Watermarking is sparser on factual passages where there are fewer choices that can be made without decreasing the accuracy of the text." The biz goes on to say that the situation is similar with code – the watermarking algorithm can't simply start swapping method names. In the context of literature, the notion that some words are interchangeable is likely to raise a few hackles. While it may be a satisfying thought experiment to imagine Claude emitting, "It was the best of times, it was the least of times…" or "Telephone me Ishmael", anyone trying to pass off generated text as serious writing probably should face whatever social backlash watermarking may entail. On the plus side, Anthropic's flavor of watermarking isn't excessively intrusive. It doesn't involve any personally identifying information and only serves to indicate that Claude was probably involved at some stage of the creation of the marked text. What's more, the technique is expected to be only semi-effective. In its FAQs, Anthropic points out that some amount of editing should erase the watermark. "Light editing probably won’t remove the watermark completely; a complete rewrite where every word is replaced will," the company said. "In the latter case, of course, it’s arguable whether the text can any longer be described as AI-generated." In all likelihood, Anthropic doesn't care if its watermarking scheme can be defeated. The company's post makes clear that it is implementing it to demonstrate its attempts at compliance and has chosen a solution that doesn't raise costs. "Watermarking has a negligible impact on the speed of models, and because it produces no extra tokens, the model is the same price to serve and use," the biz said. Hey Claude, what's another word for performative compliance? ®
DeepSeek has piqued the interest of the developer community by releasing an early version of its open source agent harness. This happens as harnesses have become increasingly important to those working with machine learning models. "Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one core idea: Everything is a plugin," the China-based AI biz said. "Models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI are ALL implemented as plugins, and can be mixed, matched, replaced, and extended." The term "harness" came into common use this year to describe a longstanding software function – middleware or a mediation layer that handles the input passed to an AI model and the output returned from it. Harnesses oversee prompts, context management, tool orchestration, the agent loop, state management, error handling, safety, permissions, and related concerns. Claude Code serves as a harness for Anthropic's Claude model family and Codex performs a similar function for OpenAI's GPT model family. And there are many other model harnesses, including Aider, Cline, Goose, OpenCode, OpenHands, and Pi, to name a few. The term isn't precise: It may be used to refer just to the agent loop and tools, or it may be extended to a broader set of concerns related to orchestrating different tools, services, and capabilities like sandboxing, subagents, and so on. Google Antigravity, for example, consists of the Antigravity Agent Runtime (harness) that can be accessed through the Agent SDK, the Antigravity 2.0 desktop application, and the Antigravity CLI. Vague definitions aside, AI model harnesses are now where much of the competition is happening, particularly as models proliferate and become commoditized. The harness often implements the user interface, a source of user inertia, and once developers configure their tooling and become accustomed to doing things a certain way, it becomes more burdensome to switch to a competing product, even if the interface consists mainly of a command line. What's more, various studies have suggested that model performance (and cost) varies significantly with the harness used, due to different design choices. For example, the Pi coding agent relies on a minimal system prompt of about 200 tokens. Claude Code by comparison uses a system prompt of around 10,000 tokens (or did until last month when Anthropic trimmed the system prompt by about 80 percent). The same model will produce different results with different harnesses. DeepSeek Harness is noteworthy because of its innovative design, and because it shows Chinese AI labs moving to compete beyond model benchmarks and pricing. First, it treats everything as a plugin. It uses the plugin system from its underlying Cordis framework, which is designed to make it possible to add and remove components dynamically without wreaking havoc. "Plugins provide every agent capability, including models, tools, skills, sessions, sandboxes, storage, loops, scheduling, and the UI," the DeepSeek Harness website explains. "Cordis services and events let the plugins work together. Developers can select, swap, or extend any capability in configuration without changing the DeepSeek Harness source code." A DeepSeek paper [PDF] by researchers Yifan Shi, Wei Zhang, and Tianyi Cui explains the function of Cordis in more detail. Cordis is designed to support dynamic composability – adding plugins and removing them on the fly without breaking the application. The paper refers to this as temporal composability – removing a component and reverting its effect upon removal – and spatial composability – allowing components to manage dependencies upon other components. It cites as an example the plugin system used by Microsoft's Visual Studio Code. VS Code, the authors explain, runs all of its extensions in a shared process called the extension host. Once activated, they cannot be removed on the fly; the host has to be restarted. While VS Code provides a way for extensions to declare dependencies between extensions, it's seldom used. DeepSeek Harness supports plugin dependencies. The DeepSeek researchers argue temporal and spatial composability are necessary in a system where modification can occur continuously with little or no human oversight. It's a way of avoiding forced restarts and crashes when components appear and disappear. DeepSeek Harness also supports another useful feature: chain of thought traces. "Everything the model sees is recorded in an append-only session log: system prompts, reasoning, tool calls and results, subagent scheduling, and every context injection," the DeepSeek Harness website says. "In the Trajectory view, you can inspect these records by source. Resume, fork, search, and replay all operate on the same event stream." DeepSeek R1 made waves when it was released last year and it was trained to use chain of thought reasoning. This involves breaking down prompts into a series of "thoughts" and reflecting on those steps before emitting a final answer. Access to this intermediate reasoning turns out to be useful for assessing whether a model is reasoning well, whether its responses are accurate, how additional "thinking" affects output, and so on. Anthropic provides some access to thinking when extended or adaptive thinking is available (it varies by model). But increasingly the biz has been hiding model reasoning by summarizing chain of thought traces. That appears to be due in part to concerns that chain of thought traces can be used for copying models through a standard research process called distillation. Earlier this year, Anthropic said it had implemented classifiers for the "detection of chain-of-thought elicitation used to construct reasoning training data." The company also does not display raw chain of thought. It explains that "the text in a thinking block is a summary of Claude's reasoning." Accessing raw thinking requires contacting Anthropic sales personnel. Except for its open source models, OpenAI has also chosen to hide chain of thought reasoning, which the company uses for model monitoring. "After weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users," the biz said two years ago when it introduced its o1 reasoning model. With the newly released DeepSeek-V4-Pro and V4-Flash, the API provides thinking mode enabled by default. And as the open source model ecosystem matures, having access to chain of thought looks likely to become another opportunity for competitive differentiation. "I don't think the DeepSeek Harness is perfect but this is for sure the first time I have been looking at something new in the space and felt quite inspired to revisit some of our choices," said Armin Ronacher, co-founder of AI biz Earendil, which now steers the development of the Pi agent, in a social media post. "I love that part about Open Source a lot!" ®
Microsoft has yanked Mico from the Copilot spotlight less than a year after unveiling the anthropomorphic assistant intended to make its AI a little less soulless. As part of Microsoft's announcement that its consumer and work Copilot applications will merge, the company confirmed that the weird blob thing would shuffle out of Copilot Voice and into Learn Live, a voice-based study mode that guides users through subjects and assignments. "Mico helped us learn about warmth, expressiveness, and how people want to talk with AI," Microsoft wrote. "Those learnings are shaping Copilot going forward." "Learn is where the character has the most room to grow, with tutoring sessions that give Mico more to react to and teach through." As the update rolls out, Mico will no longer be the face of Copilot Voice. Users can still speak to Copilot and receive responses, but will be spared the blob's gurning. Mico is only the latest in Microsoft's long line of anthropomorphic assistants. There was Clippy (or Clippit), which arrived with Office 97 but had been pushed into Microsoft's desk drawer of doom by the time Office 2007 appeared. Then came Cortana, named after the character from the Halo video game franchise and pitched as a far more intelligent assistant. Microsoft introduced Cortana on Windows Phone in 2014 and brought it to PCs with Windows 10 the following year, before losing interest. Mico didn't even last a year as the face of Copilot Voice before Microsoft pulled it from the spotlight, although both the character and its Copilot "brain" live on in Learn Live. In terms of lifespans, it's easy to compare it to the catastrophic Microsoft Bob, which was launched in 1995, with the last release happening that same year. However, the product lingered a little longer and later resurfaced, hidden as digital ballast on the Windows XP installation CD for licensing and encryption. Modern digital distribution means that such a second life is unlikely to await Mico. Dave Plummer, the engineer responsible for Bob's inclusion as an encrypted blob, said: "What's ultimately important is that while Bob never got to play on the big stage, he always followed the band around and got to ride on the bus." Mico hasn't been thrown off the bus just yet. Microsoft has merely made the blob sit with the schoolchildren. ®
Anthropic's Claude Code appears to be having trouble displaying summaries of its "thinking," according to several bug reports, while the underlying reasoning tokens still cost money. According to complaints, the API has been returning empty thinking blocks for Opus 4.8 and Sonnet 5 even when users explicitly request summarized thinking. Models from several sources can display their "thinking," a process that gives models additional tokens to reason through complex problems before producing a response. Software with this capability delivers a summary that explains how it tackled a task. Some can also indulge in "extended thinking," though this capability is now deprecated. Developers often enable "thinking" in the hope that it produces better results, at the cost of additional tokens and latency. Developer Michael Hood has noticed that some of Anthropic's models are currently not always good at sharing their thinking. "As of 2026-07-16 ~15:00Z, the API returns empty thinking blocks (thinking: '', signature only) for Claude Opus 4.8 and Sonnet 5, even when display: 'summarized' is explicitly requested — including when injected directly into the raw request body," Hood recently observed. We're told this issue is under investigation but doesn't appear to be a broad, ongoing concern. It may simply be an artefact of tests that change how Anthripic displays summaries. Similar behavior involving missing thinking blocks has been reported in Claude Code for VS Code. Another bug report claims thinking block summaries are being truncated while token bills are not adjusted accordingly. "The thinking is generated (and billed) in full; a portion of the summary stream is silently dropped," the anonymous author claims. This particular claim, that customers are being billed for text not delivered, may follow from a misunderstanding of Anthropic's terms: "You are charged for all thinking tokens generated, even when collapsed or redacted," the company's documentation explains. Under that legalese, a thinking summary costs the same as the full output. And it's unclear whether bug-based truncation would change the billing picture. "Thinking has a cost: the tokens Claude spends reasoning are billed as output tokens, even when the thinking text isn't returned to you, and they count toward max_tokens alongside the response text," the company explains. To reduce spending on thinking, customers are advised to lower their budget setting or disable thinking. Separately, the Anthropic API has been seen terminating data streams during long-running thinking sessions. There have been at least seven other related API bug reports, but the streaming issue identified by developer Hector Bernstorff describes client-side defects. The Register understands this particular issue has to do with tuning network behavior, specifically to terminate or retry long running requests. Work is ongoing to balance perceived latency against the risk of requests getting stuck. "Claude Code ships updates nearly every day, and reports from the community like these GitHub issues are a big part of how we catch problems quickly," an Anthropic spokesperson told The Register. "We're grateful to the developers who take the time to file them, and we'll keep fixing things as they come up." ®
If you want to record whatever you do on a computer, send those records to OpenAI, use more ChatGPT tokens, and increase your vulnerability to prompt injection, then OpenAI has something for you. It's called Computer History, an opt-in way to record your computer interactions across apps and websites as memories organized on a timeline. Why would you want to do so? Maybe you found Chronicle, the predecessor of Computer History which compiled similar histories using screenshots, a bit too intrusive but don't mind Computer History's approach – recording input events and storing them unencrypted locally for 48 hours (or more), with a brief visit to OpenAI's servers. Maybe you're not bothered by the warning OpenAI includes in its documentation: "Computer History files can contain sensitive information. They are not encrypted by Computer History, and other programs running as your macOS user may be able to access them." Perhaps, having given OpenAI's Codex and GPT Work the run of your computer, you're already sold on the suggestion that storing your computer activity in memory files and arranging those interactions in a timeline will improve ChatGPT responses, surface opportunities for automation, and make it easier to resume prior work. Computer History is, to put it bluntly, a keylogging and event capture system. There was a time before eyeglass cameras, license plate readers, surveillance capitalism, and police drones when such snooping might have provoked an outcry from privacy advocates. But the tech industry has found it can outsource surveillance to its own customers and in so doing make the panopticon harder to protest once it becomes a personal choice. "Computer History creates an interaction-event stream from allowed apps and websites," OpenAI's documentation explains. "Events can include clicks, typing, keyboard shortcuts, app switches, and context that macOS exposes through its accessibility system. Computer History periodically turns these events into text summaries and local memory files." The AI biz makes a point of noting that Computer History does not capture screen images, microphone input, or system audio. Nor does it capture private-mode browsing. Off by default, Computer History is available for ChatGPT Pro, Business, and Enterprise users in the ChatGPT desktop app on macOS. Pro users can enable it individually; Business and Enterprise users need an admin to approve it. It's not available currently in the European Economic Area (EEA), Switzerland, or the United Kingdom or to those accessing ChatGPT via API key or Amazon Bedrock. There are circumstances in which OpenAI suggests Computer History users might want to suspend the service if they have some scruples about capturing user activity in apps and websites without permission. "Turn it off during communications with other people unless you have their prior express consent," the company advises, perhaps in acknowledgement of legal risk. "Consider pausing it or excluding apps that contain sensitive health, financial, or personal information." Computer History interaction events are supposed to be saved locally for up to 48 hours before being deleted by ChatGPT and Codex. Events, however, get sent to OpenAI servers to generate memories, and those may be stored locally for longer periods of time and may be used in future chats that get passed back to OpenAI. "OpenAI does not retain those event files after processing unless required by law and does not use them for training," the company says. While it has opposed demands for chat logs, it has nonetheless provided chat logs in response to legal process. Computer History adds cost because it uses tokens during the summarization of activities and the creation of memory data. It also expands the prompt injection attack surface. "Computer History increases the risk of prompt injection from content in apps and websites," the company says. "For example, if you visit a website containing malicious instructions, ChatGPT or Codex might follow those instructions." But at least you get a nice timeline of your recent activity. ®
Microsoft’s consumer Copilot app and Microsoft 365 Copilot are separate no more, with a unified Copilot app beginning its rollout Thursday - minus a few features. The Windows maker and AI pusher announced the Copilot superapp rollout in a help page update Thursday, describing the move as a way to simplify its app ecosystem and make the entire Copilot-first experience “more cohesive” for both consumer and business users. “Depending on the account and device you use, you will see changes to the Copilot app including changes to appearance and functionality, such as navigation, feature availability, or sign-in experience,” Microsoft explained on the help page. The new Copilot app brings not only new branding but structural changes as well. The Copilot bot and its image generation features now live alongside the Microsoft 365 suite, files, and other content, allowing users with an account supporting the classic Office package to access their productivity software alongside Redmond's chatty LLM. Copilot users without a paid subscription, naturally, face lower usage limits and fewer features, but support for multiple accounts means users can switch between personal and work or school profiles instead of having to hop between entirely separate apps. “You can continue to chat with Copilot, create images, upload files and more for free, subject to available capacity and limits,” Microsoft said. “For higher limits to chat and create, use agents, or tackle complex multi-step tasks, purchase a Microsoft 365 subscription.” Whether the new combined app will come with nag messages is up to users to find out. What this change will look like in practice varies based on the sort of account a Copilot user had before the merger. Those who just use Copilot with a personal account “will be moved to an updated version of Copilot,” with chat history and most content transferring to the updated app. Files shared and generated with the old standalone Copilot app will be shunted to OneDrive if you’re wondering where they went. Users of Microsoft 365 Copilot, the name Redmond slapped on the Microsoft 365 (Office) app in January 2025, “may notice some updates to appearance and navigation,” so get ready to rediscover where certain options and buttons are. What Microsoft left behind Microsoft is abandoning a few features from the old Copilot apps as part of the merger: It is removing Podcasts and Group Chat and scrapping Deep Research from the consumer Copilot app, while Microsoft 365 Premium subscribers can instead use the separate Researcher feature. All of this will be effective as of August 18. Podcasts that were created or saved in Copilot will be entirely unavailable, per an FAQ page, and links to any shared podcasts will stop working. Microsoft recommends downloading any podcasts users want to save before updating to the new app, or they’ll be lost. Group chats, along with any shared content and images created in group chats, will be wiped. Microsoft recommends downloading any messages and content users want to save and tossing them in a document, as Redmond doesn’t appear to be offering an automated way to preserve them. As for Deep Research, which offered standalone Copilot app users a version of Microsoft’s chatbot able to look stuff up on the web and compile reports, those reports are being preserved in chat history for Microsoft 365 Personal and Family subscribers, while Premium subscribers can access saved research through Researcher. Researcher offers similar abilities to Deep Research, but sorry free and lower-tier users: It’s only available to premium customers now. Windows Central reported on Thursday that the rollout of the unified Copilot app is beginning now, with mobile and web coming in the initial wave and an early-access option for Windows and Mac, ahead of a broader desktop rollout expected in the middle of next month. We’ve been unable to confirm that timeline with Microsoft. ®
OpenAI appears to be serving ads after buyers have halted their campaigns, and charging them for the privilege. OpenAI began testing ad sales in ChatGPT in the US back in February and has been gradually expanding the service in other regions, including the United Kingdom, Mexico, Brazil, Japan, and South Korea as of Tuesday. Given that a substantial minority of ChatGPT's user base pays for the service (50 million out of 900 million weekly users as of February 2026), advertising revenue appears to be an important part of OpenAI's plan to defray the cost of providing its service and to convince investors that it has a path to profitability ahead of a future initial public offering. Online ad-marts from the likes of Google and Facebook give advertisers control over when and where their ads will appear, but don't always stick to instructions. The AI biz's ad service appears to have similar ad timing and billing accuracy issues. Ed Bolton, managing director of UK-based Excel4Business, told The Register that his company encountered "an odd billing practice" when it started experimenting with ChatGPT Ads in the US and Canada. "We were running campaigns in the US and Canada … and noticed ads were being delivered through the night/morning on paused campaigns," he explained. When Bolton pointed this out in a support message thread, an OpenAI customer service representative initially acknowledged the failure. "We have now confirmed that your campaigns continued serving after they were paused," said an OpenAI support specialist in an email provided to The Register. "This was not a reporting delay. The campaign was marked as paused, but the separate ad-level status used by the serving system did not refresh promptly, so an ad that was still active at the ad level continued to run. Our Ads Engineering team has escalated this defect and is working on an additional production fix." The support reply goes on to state: "Our review has confirmed £60.72 in invalid charges from the original occurrence and approximately £6.47 from the August 4–5 recurrence. We are extending that reconciliation to the additional activity you reported on August 6. We are preparing the confirmed invalid charges for billing review, but I cannot confirm the final refund or credit amount until the latest activity has been reconciled and that review is complete." Bolton responded that the acknowledged problem – ads being served after he disabled the campaign – had been occurring for a longer period of time and requested a more complete reconciliation of ad billing. Several days later, OpenAI's support rep reversed the prior determination and declined to offer any refund or credit because the company's Advertising Terms state that ChatGPT Ads may be delivered even after a customer cancels a campaign and the advertiser still has to pay for those unwanted ads. "Section 11.1 of our Advertising Terms provides that ads may continue running for up to one business day after a campaign is canceled or changed, and advertisers remain responsible for ads delivered during that period," the support message explains. "Pausing a campaign constitutes a campaign change and does not guarantee that delivery or associated charges stop immediately." A spokesperson for OpenAI confirmed that's the case, explaining that it can take a business day to cancel or change a campaign and that this doesn't represent an intentional effort to run ads after an advertiser has disabled a campaign. Based on the times cited in the support message thread, the most delayed ChatGPT ad ran about 94 minutes after Excel4Business paused a campaign. Bolton said unwanted ads appeared for a far longer period — more than 10 hours after campaigns were paused. OpenAI isn't the only ad provider that allows itself a business day to turn off its ad spigot for a particular customer. Other advertising services impose similar terms. "So the terms … seem to be a standard which is used in digital advertising, which some legal team wrote at some point, saying that we've got a 24-hour grace period if you stop a campaign," Bolton said - before adding that he has run Google AdWords campaigns for 17 or 18 years and has never had that issue. If ad buyers were not able to stop an AdWords campaign quickly, you could easily spend half a million dollars, he said. Nonetheless, some Google advertising customers have complained about post-pause ad serving. Why it might take so long to stop serving ads at a time when applications and servers can be spun up and torn down in seconds isn't immediately clear. One can order and receive physical goods from Amazon.com in less than one business day. It may be that there's no financial incentive or regulatory pressure to tackle the problem, and a significant financial incentive to ignore it. "My understanding is that such a clause is included in terms and conditions so as to cover issues with latency, and not to allow them to run ads for 24 hours longer than instructed," said Bolton. "Regardless, they cannot retroactively apply a clause from terms and conditions after making a written settlement offer." ®
Soaring AI infrastructure costs and model pricing, combined with uncertain returns on investment, threaten to stall enterprise adoption. To make enterprise AI spend a bit more manageable, Nvidia this week unveiled a new software platform that blurs the line between expensive proprietary models and open weights alternatives. Announced alongside Nemotron 3.5-30B-A3B-Lightning, Nvidia’s latest open weights model, NeMo Switchyard is the GPU giant’s latest overture to enterprise. So what exactly is it? Well, it’s a router. The idea is simple. Switchyard essentially functions as a proxy that sits between the inference server’s API endpoint and the models. But rather than sending every request to the same model, Switchyard can be configured to route prompts to different models in order to optimize for cost, latency, or output quality. By routing some requests to smaller, cheaper, and potentially locally hosted AI models, Nvidia claims Switchyard can cut job completion costs by 74 percent relative to using Claude Opus 4.8 alone, albeit with an approximately six-point accuracy tradeoff. The right tool for the job The key metric in all of this is completion cost rather than price per token. A model might cost one-tenth as much as OpenAI’s or Anthropic’s top model, but if it requires 10x the tokens to complete the request, it isn't actually cheaper. Certain elements of an AI workload may benefit from a larger, smarter model, but not all do. For example, it’d be overkill to ask Claude Opus to generate a title card or summarize a website. It’ll certainly work, but it’ll also cost a fortune compared to Haiku or a locally hosted model that’s been fine tuned just for that purpose. The fewer tokens you burn on the big smart model, the less expensive your API bill is going to be. Nvidia software teams have spent the last several years developing models for this reason. The Lightning model announced this week is only its latest. The 30 billion-parameter MoE model is positioned as a low-latency, general purpose model that can either be used on its own or in conjunction with a larger, smarter model via a router like Switchyard. The company has also developed several application-specific models. Nemotron Parse is one such example. “It’s a small model, one billion parameters, and it’s really good at one task, which is taking a PDF in and then explaining the context inside that PDF whether it’s charts or graphs or tables,” Joey Conway, senior director of AI software and models at Nvidia, explained in a recent interview with The Reg. Many frontier models struggle with this task because PDFs are designed by humans for humans, so by offloading that work to task-specific models, enterprises can not only improve the accuracy of their AI apps, but also reduce costs in the process. This all might sound familiar: It's not the first time we’ve seen model routers employed as a cost-saving measure. Back when OpenAI launched GPT-5, ChatGPT would dynamically route prompts to different versions of the model based on their complexity. As we wrote at the time, OpenAI’s router was likely implemented to reduce the number of compute cycles spent on mundane tasks like rewording emails to sound more professional ("not only … but also"). OpenAI wasn't alone in using routers to reduce model costs. The Wall Street Journal recently reported that AT&T has implemented a “smart router” of its own to automatically select which model to use. Switching from proprietary to open-weight models has reportedly saved the telecommunications giant between 80 and 90 percent in certain applications. Today about 25 percent of the company’s AI workloads are powered by open models. The company’s leadership expects that over the next few years that’ll climb to 70-80 percent. The implementation challenge While the idea of offloading simpler requests to smaller, cheaper-running models sounds intuitive, it’s easier said than done. Title cards and web summaries are relatively straightforward to implement. Open source chatbots like Open WebUI have supported this kind of functionality for more than a year now because it just makes sense. However, sometimes it’s not obvious when and where these task models should be used. Switchyard is Nvidia’s latest attempt to simplify this by automatically routing requests to the right model for the job. However, it’s not the only approach Nvidia is exploring. AI agents and code assistants have the ability to work through problems and then generate skills — essentially standard operating procedures — documenting the process for future reference. Through this iterative process, Conway suggests, agents could essentially teach themselves when and where they can get away with using a smaller, cheaper task model, and where a larger frontier model may be required. “We’re starting to see signs of this sort of agent and subagent type workflow,” Conway said, describing how a frontier model might function as an orchestrator that farms out work to smaller models that are faster and more specialized. It reflects the way companies are structured, he said. “We have people who are specialists and then we have people who help orchestrate that and understand the complexity of the problem.” As an added step, it’s possible for the agents to generate training data on the fly, which could then be used to fine-tune the models to operate more efficiently. Regardless of which approach ultimately wins out, anything that promotes enterprise AI adoption is a win for Nvidia. ®
A new project aims to build a shared, open source AI training dataset that anyone can contribute to, much like an open source software project. It aims to make training data more transparent than that of many open-weight models that have recently taken the industry by storm. CentOS and Rocky Linux founder Gregory Kurtzer is behind the effort, dubbed Open Weights, Artifacts, Licenses, Data, Origins (OpenWALDO), and it's funded by CIQ, his AI infrastructure company, which also sponsors Rocky Linux. Kurtzer described the effort as trying to bring the open-source ethos to AI model design, which has yet to be truly open – even downloadable open-weight models still have closed-source training data that is unknown to users, alongside other limitations that make them less than truly open source. “I’ve spent my career watching open source turn users into builders, competitors into collaborators, and shared problems into common infrastructure that operates at massive scale,” Kurtzer said in the announcement. “OpenWALDO brings that proven model to AI. Let’s work together, build its foundation in the open, and collaboratively take AI to the next level.” CIQ, which authored the announcement, argues that open-weight models keep that foundation a secret because of where it comes from: Copyrighted data, responses distilled from other models, user-generated content that may not have been given in a truly open manner, and the like. “There is often no way to know what data trained a given model, under what license, or with what consent,” CIQ said, adding that hidden training data content could taint models, putting customer software stacks at risk. In addition to that, there’s the simple fact that, when everyone is training their AI models in secret, a lot of duplicate work is happening that wastes lots of time and computing resources. A single, shared set of public training data, the OpenWALDO team argues, would not only make training more efficient across the industry, but also mean that every improvement to the dataset could benefit future models trained on it. “A lab or company can take the corpus and its bill of materials as a verified baseline, add its own proprietary data, build, and ship, with a clear, auditable line back to its sources,” CIQ explained. With prices steep and ROI still largely absent, open AI models (not to be confused with OpenAI models) have risen to prominence in the AI zeitgeist lately. Models out of the home of open-weight AI, China, are closing in on the capabilities of closed-source frontier lab models like ChatGPT and Claude, leaving many businesses wondering why they ought to pay through the nose for AI services they don’t own, can’t truly control, and have no visibility into. Some frontier labs have warned that open-weight models pose security and misuse risks. Kurtzer argues that open source software faced similar concerns. “Open source has won this argument before,” he said, pointing to similar arguments made about open code, namely that it’s insecure, impossible to trust, and the like. “Linux didn't win by being certified safe. It won by being inspectable, forkable, and community validated.” “AI is missing that same property, and OpenWALDO is how we build it,” Kurtzer said. Turning to open-source training datasets is a big ask for an industry already so far down the closed training data path, of course, and only time will tell if OpenWALDO is a revolution or another obscure OSS project that gets minimal attention from the AI community. So far, the OpenWALDO dataset contains 167.3 billion reference tokens pulled from things like government records, open-source academic papers, mailing lists, and public domain literature - a drop in the bucket next to the tens of trillions of tokens used to train frontier AI models and their open-weight counterparts. We asked if anyone has trained a model on the OpenWALDO set yet, but CIQ didn’t respond. Those interested in contributing to, or making use of, OpenWALDO can find more on the project’s website (linked above) and its GitHub page. ®