OpenAI 发布 Astra:主攻网络安全与软件工程,但可解释性引发争议 OpenAI Launches Astra for Cybersecurity and Software Engineering, Raising Interpretability Concerns
据 TechCrunch AI 转述,OpenAI 称 Astra 是其最新、最强的 AI 模型,具备电脑与浏览器操作能力,并计划分阶段向网络安全客户、付费用户和 API 开放。其 opaque recurrence 推理技术也让外界担心模型决策过程更难监控。 According to TechCrunch AI, OpenAI says Astra is its latest and most capable AI model, with computer and browser-use abilities. It is slated for staged access through a cybersecurity program, paid plans, and the API, while its opaque recurrence approach has raised concerns about monitoring model reasoning.
OpenAI 表示,Astra 是公司最新、也是目前最强的 AI 模型,重点提升电脑和浏览器使用能力。公司将其定位为更快、更准确且更安全的模型,但现有证据主要来自 OpenAI 的发布说法。
发布与可用范围
- 首批用户:Daybreak 网络安全项目的客户。
- 后续开放:OpenAI 计划在一周内向 Pro、Plus、Enterprise 和 Business 付费用户开放。
- 开发者入口:模型也将加入 API。
主要应用方向
Astra 的重点应用包括网络安全和软件工程。OpenAI 称,模型在安全测试中能够帮助防守方发现并修补 zero-day 漏洞,也就是尚未被公开修复的安全缺口。公司还称 Astra 是其最强的软件工程模型,在查找 bug、执行终端任务和回答代码库问题等测试中取得较高成绩。
OpenAI 将 Astra 描述为目前最聪明、也最符合目标的模型;这些判断属于公司及其管理层的公开表述。
争议集中在推理可观察性
Astra 使用名为 opaque recurrence 的技术。根据现有材料,这种方法可能让研究人员更难观察模型如何逐步形成结论,从而削弱通过 chain of thought 线索检查决策依据的能力。OpenAI 表示已加入新的安全措施并进行了多项安全基准测试,但也承认,模型能力越强,监控其内部推理过程可能越困难。
AGI 表述仍未形成明确节点
OpenAI 总裁 Greg Brockman 将 Astra 形容为公司当前最聪明、最对齐目标的模型,并表示它可能改变人们交给 AI 的工作范围。当被问及这是否意味着 AGI 已经到来时,他没有将其确认成合同层面的节点,并表示 AGI 更像使命或精神概念;他个人认为 OpenAI 已达到这一阶段。
OpenAI says Astra is its latest and most capable AI model, with improved computer and browser-use abilities. The company positions it as faster, more accurate, and safer, although the available evidence is primarily based on OpenAI’s own launch claims.
Launch and Access
- Initial users: Customers of the Daybreak cybersecurity program.
- Next phase: OpenAI plans to make Astra available to Pro, Plus, Enterprise, and Business paid users within a week.
- Developer access: The model is also expected to be offered through the API.
Primary Use Cases
Astra is being promoted for cybersecurity and software engineering. OpenAI says it performed in security testing aimed at helping defenders identify and patch zero-day vulnerabilities, meaning security flaws without a publicly available fix. The company also calls Astra its strongest software engineering model and says it scored highly on bug finding, terminal tasks, and codebase question answering.
OpenAI describes Astra as its smartest and most aligned model to date; those characterizations are company and executive claims in the available evidence.
Interpretability Is the Main Concern
Astra uses a technique called opaque recurrence. Based on the available material, this approach may make it harder for researchers to observe how the model reaches conclusions, reducing the usefulness of chain-of-thought signals for auditing decisions. OpenAI says it added new safeguards and ran multiple safety benchmarks, while acknowledging that stronger models can be harder to monitor because they may complete more difficult tasks with less or no visible reasoning text.
AGI Claims Remain Ambiguous
OpenAI president Greg Brockman described Astra as the company’s smartest and most aligned model, saying it could change the range of work people delegate to AI. When asked whether this meant AGI had arrived, he did not identify it as a contractual milestone and said AGI is better understood as a mission or spiritual concept. He personally believes OpenAI has reached that stage.
来源
- TechCrunch AI · 09-04 02:01