GPT-6 Astra: OpenAI’s New AI Model Pushes Toward a New Era of Autonomous Work
OpenAI has unveiled GPT-6 Astra, describing it as its most capable and aligned AI model so far. The company says Astra represents a major step beyond conventional chatbot capabilities, combining advanced reasoning with computer use, software engineering, scientific research and professional knowledge work.
Astra is designed not merely to answer questions but to perform multi-step tasks on a computer. OpenAI says it can fill online forms, update CRM records, organize calendars, conduct web research, work with documents and spreadsheets, create websites and troubleshoot software. It can also operate specialized scientific applications and inspect research data.
One of the biggest changes is its emphasis on autonomous computer use. In OpenAI’s testing, Astra scored 72.6% on OSWorld 2.0, compared with 65.7% for GPT-5.6 Sol. OpenAI also says Astra completed comparable computer-use tasks in roughly 40 minutes versus about 75 minutes for Sol in its latency simulations.
The model is also aimed at professional work. OpenAI says Astra can produce documents, presentations, spreadsheets and analyses while following existing templates and organizational styles. The company says the model is better at understanding when an instruction leaves room for interpretation and deciding when to make a reasonable assumption versus when to ask the user for clarification.
Coding is another major focus. OpenAI calls Astra its strongest software-engineering model to date, with improvements in agentic coding, testing and long-running development tasks. A new Codex capability allows Astra to preserve and retrieve information across context windows instead of relying entirely on compressed summaries of earlier work.
Its reported benchmark results are particularly striking. OpenAI says Astra achieved 96% on GPQA Diamond, a graduate-level scientific reasoning benchmark, and 57.9% on Terminal-Bench 4.0. On BenchCAD, which tests reconstruction of 3D objects through CAD code, Astra reportedly reached a 95.9% geometric-overlap score.
The model also represents a significant development in scientific AI. OpenAI says Astra has helped tackle difficult mathematical and scientific problems and can combine reasoning with computer interaction to work inside specialized research software. The company’s claims come amid a wider push by frontier AI labs to move from AI that explains research to AI that actively participates in the research process.
Cybersecurity is simultaneously becoming one of the most sensitive areas. OpenAI says Astra meets the Critical threshold for cybersecurity under its Preparedness Framework because of its ability to identify and develop sophisticated exploits. That capability could help defenders discover vulnerabilities, but it also increases the risks associated with deploying increasingly autonomous AI systems.
The model’s availability is expanding across the technology ecosystem. OpenAI says GPT-6 Astra is being rolled out to ChatGPT Plus, Pro, Business and Enterprise users and is available through the OpenAI API, Microsoft Azure and Amazon Bedrock. The API model uses the identifier gpt-6-astra, with standard pricing listed at $10 per million input tokens and $50 per million output tokens.
Astra is already finding its way into specialized industries. OpenAI’s new financial-services offering, launched with partners including Morgan Stanley and Evercore, uses GPT-6 Astra for financial research, modelling and client materials while integrating professional data sources such as LSEG and PitchBook.
But the launch has not been without controversy. Some users have recently claimed that Astra’s performance changed after launch, with complaints that the model became less capable or behaved differently. These reports are user observations rather than established evidence of an intentional downgrade, but they have revived a familiar debate about how frontier models behave after deployment.
The larger question is whether GPT-6 Astra marks the point where AI moves decisively from a chatbot into an autonomous digital worker. Its combination of reasoning, computer control, coding and professional workflows suggests that the competitive battle is increasingly shifting from who can produce the best answer to who can safely complete the most complicated real-world task.
For businesses, developers and researchers, that distinction could be significant. If Astra’s capabilities translate reliably from benchmarks into everyday work, AI systems may increasingly handle entire workflows rather than individual steps—potentially changing how software is developed, research is conducted and professional services are delivered.
