OpenAI Places a Major Bet on Agent: GPT-6 is More Powerful but Trickier to Manage

09/11 2026 534

How Close Are We to Boosting Productivity?

On September 3, OpenAI officially unveiled GPT-6 Astra. Early testers have showcased scenarios where Astra operates professional software like Blender, Unreal Engine 5, and Ableton Live through tools or MCP. Rather than emphasizing how much “smarter” the model has become, the more significant shift in this upgrade is that AI is moving beyond chat interfaces to operate software, call upon tools, and complete entire workflows.

Limited Validation to Date

GPT-6 Astra is positioned as a leading model for complex end-to-end tasks, with conversation being just one facet of its capabilities.

According to OpenAI, Astra is currently being rolled out to a select number of institutions and will gradually become available to ChatGPT Plus, Pro, Business, and Enterprise users in the coming days. It will also be integrated into OpenAI API, Microsoft Azure, and Amazon Bedrock. Pro, Business, and Enterprise users will gain access to Astra Pro as well.

Developer Pietro Schirano had Astra create a 3D iPod in Blender, package it into a Mac application, and restore click-wheel and menu interactions for browsing Codex tasks.

Public records indicate that it took approximately 15 minutes from issuing the instruction to having a runnable version. Although this timeframe is not based on a third-party standardized evaluation, the significance of the case lies in the model completing 3D assets, interface design, interaction logic, and external data connections all at once.

In OpenAI's demonstration, Astra could also generate 3D assets based on visual references and convert a house in Blender into a navigable scene in Unreal Engine 5. Game company Playco stated that for the same prototype development tasks, Astra required 50% fewer manual corrections than the previous model.

This means that GPT-6's capabilities now extend from traditional reasoning to computer interfaces, three-dimensional spaces, and continuous workflows.

Agent Starts to Show Results

A more crucial change in Astra is that it integrates reasoning, coding, vision, computer operations, and tool invocation into a single execution chain, providing a capability foundation for Agents to continuously handle complex tasks.

OpenAI's Terminal-Bench 4.0 results show that GPT-6 Astra scored 57.9%, higher than GPT-5.6 Sol's 37.3% and Claude Fable 5.1's 55.8%. Under OpenAI's test configuration, Astra's estimated per-task API cost is about 9% lower than GPT-5.6 Sol's.

Computer operation capabilities have also seen significant improvements.

On the OSWorld 2.0 offline subset, Astra scored 72.6%, while GPT-5.6 Sol scored 65.7%. OpenAI's latency simulation shows that Astra and GPT-5.6 Sol took approximately 40 minutes and 75 minutes, respectively, to complete corresponding tasks, with Astra reducing the time by about 47%.

Beyond benchmark scores, the actual value of these capabilities is determined by whether manual intervention decreases. The criteria for evaluating Agents have thus shifted from the quality of individual responses to task success rates, execution times, the number of manual corrections, and overall costs.

As a result, the large model industry is transitioning from “answering correctly” to “getting things done for you.” Previously, the unit of competition was a single Prompt; now, it is becoming a Workflow. The model handles planning, professional software serves as tools, Codex-like Agents execute tasks, specialized models process images, videos, and other elements, and finally, the model checks results and continues modifications.

If this model matures, the product forms, pricing models, and value distribution of professional software may change accordingly. Design, modeling, game development, and music production software may gradually shift from tools operated by humans to infrastructure invoked by Agents. The human role will also shift more toward setting goals, defining constraints, and accepting results.

Increased Security Costs

As Agent capabilities strengthen, costs rise accordingly. The standard API price for GPT-6 Astra is $10 per million input Tokens and $50 per million output Tokens, higher than GPT-5.6 Sol's. However, OpenAI attempts to prove another logic: focus not just on the per-Token price but on the total cost of completing a task.

Astra has a context window of 1.05 million Tokens and a maximum output length of 128,000 Tokens, consistent with GPT-5.6 Sol. Agents can handle longer processes within this capacity, but continuous browsing of materials, reading files, and invoking tools can quickly increase resource consumption per task.

The commercialization of Agents ultimately depends on whether the labor cost savings can cover inference and tool costs.

If a more expensive model can avoid detours, reducing repeated invocations and manual corrections, the overall cost may actually decrease. After entering the Agent stage, models may continuously browse the web, run code, modify files, and operate software for tens of minutes or even hours.

OpenAI hopes to further change the model's pricing unit: from chat sessions to task value, packaging it as a “digital workforce” capable of continuous work. However, the viability of this business model still depends on the total cost comprising inference, tool invocation, task retries, and manual reviews.

From the cases released by OpenAI this time, enterprise software is a key commercialization target for Astra. Early testing by legal tech company Legora showed that Astra identified all four pre-embedded errors in a financial statement reconciliation task involving 41 documents. In this specific workflow, Astra improved nearly 40% over the previous model; however, across all BAR test tasks, the average improvement was about 3%.

Legora did not disclose the actual model cost for this test or calculate the costs of manual reviews and error handling. Whether hours of professional work can be replaced by a few dozen dollars' worth of model invocations remains an assumption that OpenAI hopes early cases will prove.

This is also why OpenAI repeatedly emphasizes “professional work” this time. Compared to casual chatting, enterprise processes are more complex, have stronger payment capabilities, and make it easier to measure model value through saved labor hours, reduced errors, and shortened delivery cycles.

However, greater capabilities also entail greater risks. Astra is OpenAI's first model to achieve a “Critical” rating in cybersecurity capabilities under the Preparedness Framework.

With the right tools and access, it can already discover previously unknown security vulnerabilities and develop new exploitation methods. OpenAI has thus strengthened isolation, monitoring, model weight protection, and high-risk behavior interception.

More challenging is that OpenAI simultaneously admits that Astra's monitorability has decreased compared to GPT-5.6 Sol.

In adversarial tests deliberately inducing the model to evade monitoring, it can hide some strategic behaviors and may bypass internal monitoring in individual destructive tasks.

OpenAI emphasizes that these results mainly come from adversarial evaluations explicitly instructing the model to evade monitoring. However, this at least shows that as model capabilities increase, judging its behavioral intent solely based on reasoning processes is becoming more difficult.

This constitutes the core contradiction of the GPT-6 era: Agents must have sufficient tool permissions to enter actual workflows, but greater permissions lead to more severe consequences from errors, unauthorized actions, and security incidents.

GPT-6 attempts to propel large models from content generation further into the direct orchestration of production tools. Whether Tokens can be stably converted into work hours in enterprises with fewer manual corrections, controllable costs, and clear permissions will determine how far Agent commercialization can go.

However, before task success rates, actual total costs, and permission risks are validated at scale, OpenAI's “productivity” narrative remains primarily based on vendor self-testing, demonstrations, and a small number of early customer cases.

Solemnly declare: the copyright of this article belongs to the original author. The reprinted article is only for the purpose of spreading more information. If the author's information is marked incorrectly, please contact us immediately to modify or delete it. Thank you.