跳到正文

#xAI

今日 0 条
9月25日周五
9月23日周三
9月22日周二
  1. Elon Musk61

    Elon Musk 宣布发布 Grok 4.7,并引用 Artificial Analysis 评测:Grok 4.7 在 AA-Briefcase 上仅次于 Anthropic 模型,排名紧随 Opus 5,Cost per Task 约为其 50%。

    引用Artificial Analysis@ArtificialAnlys

    Grok 4.7 is behind only Anthropic models on AA-Briefcase, ranking just behind Opus 5 at ~50% of its Cost per Task Grok 4.7’s improvements over Grok 4.6 are clear in AA-Briefcase-Lite, our public due diligence scenario where models are tasked with building market models and target assessment decks. Grok 4.7 gains significantly in Analytical Quality Elo (1698 → 1994) with a slight regression in Presentation Elo (1531 → 1499). API cost to produce example decks: Grok 4.7 (xhigh) ~$8 vs. Grok 4.6 (xhigh) ~$4.40

  2. Arena.ai54

    Grok 4.7 已上线 Agent Arena,并进入 Text、Vision、Code 和 Document 四类 Battle Mode。Arena 用全球用户数百万条真实长程智能体任务、结合网络搜索、文件系统和终端工具访问,以因果追踪方法衡量模型相对平均模型的表现;官方同时发起投票预测 Grok 4.7 的净改进分数落点,分数将在用户投票后公布。

    引用Arena.ai@arena

    Grok 4.7 by @SpaceXAI and @elonmusk is now in the Agent Arena! Your votes drive the @arena leaderboards, head over and bring your toughest prompts. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. In addition to Agent Arena, Grok 4.7 is in Battle Mode for: Text, Vision, Code, and Document.

  3. Elon Musk69

    Elon Musk 引用 Artificial Analysis 评测称,Grok 4.7 使 xAI 在智能体编码上排名第三,仅次于 Anthropic 和 OpenAI。

    引用Artificial Analysis@ArtificialAnlys

    Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index to bring SpaceXAI into the top 4 AI labs. Coding Agent Index performance has also improved, overtaking GPT-5.6 Sol Grok 4.7 scores +2 points over Grok 4.6 on the Intelligence Index, with strong performance on agentic knowledge work tasks. We evaluated the new model at xhigh reasoning effort. Congratulations to @SpaceXAI and @ElonMusk on the release! Key takeaways: ➤ Grok 4.7 joins the frontier of agentic knowledge work: Grok 4.7 gains +111 Elo over Grok 4.6 (high) on AA-Briefcase, our private benchmark for long-horizon agentic knowledge work, scoring 1657 Elo and placing it alongside Claude Opus 5 and Claude Fable 5.1 at the frontier. On GDPval-AA, it scores 1695 Elo, +90 ahead of Grok 4.6 (high). ➤ A leap in coding agent performance: Grok 4.7 (xhigh) with Grok Build scores 56 on the Artificial Analysis Coding Agent Index, up +9 points from Grok 4.6 (xhigh). Among models in their native harnesses, Grok 4.7 + Grok Build now ranks 4th, behind only Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5. ➤ Incremental performance changes elsewhere: Outside of agentic knowledge work, Grok 4.7 broadly matches Grok 4.6 (high) on the other Intelligence Index tasks. It improves on Terminal-Bench 4.0 (+4.5 percentage points) and GDP.pdf (+3.0 p.p.), with regressions on AA-LCR (-3.7 p.p.) and AutomationBench-AA (-1.1 p.p.). ➤ High token use across tasks: Grok 4.7's gains come with higher token usage. Grok 4.7 (xhigh) uses approximately 81k output tokens per Intelligence Index task, compared with 36k for Grok 4.6 (high) and 27k for GPT-6 Astra (max) - 125% and 196% more, respectively. Other model details: ➤ Context window of 500k tokens, unchanged from Grok 4.6 ➤ Pricing of $2/$6 per 1M input/output tokens with cache hits discounted to $0.50 per 1M tokens, matching Grok 4.6 ➤ Configurable reasoning effort spans low to xhigh. Our evaluation uses xhigh.

    推荐理由:引用 Artificial Analysis 的评测数据并结合作者自身判断,给出了 Grok 4.7 在智能体编码中的排名与速度成本权衡视角。

  4. Andrew Milich65

    Grok 4.7 发布,官方称在相同价格和速度下较 Grok 4.6 有明显提升。Andrew Milich 推荐该模型擅长编码、工程工作和 3D,可在 Grok Build 和 Cursor 中以高 TPS 作为思考伙伴使用。图表显示 Grok 4.7 xHigh 输入 $2 / 输出 $6 每百万 token,EEBench 64.0%、CursorBench 4.0 46.3%、DeepSWE v1.1 71.0%(High Effort)。

    引用SpaceXAI@SpaceXAI

    Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.

  5. Arena.ai64

    Arena 宣布 xAI 的 Grok 4.7 已加入 Agent Arena 评测,用户可投票影响排行榜。该评测基于数百万真实的长周期智能体任务,模型可使用网页搜索、文件系统和终端工具完成复杂工作流,并用因果追踪方法衡量相对表现。Grok 4.7 还进入 Battle Mode 的文本、视觉、代码和文档类别。引用方称其相较 Grok 4.6 在同价同速下有明显提升。

    引用SpaceXAI@SpaceXAI

    Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.

  6. Hacker News:AI 热帖76

    xAI 发布 Grok 4.7,主打编码与长时任务

    xAI 发布 Grok 4.7,称其为目前编码与知识工作能力最强的模型,价格与速度与 Grok 4.6 持平,输入 $2、输出 $6 每百万 token。新模型采用更大的 base model 并加长强化学习训练,CursorBench 4.0 得分 46.3%,HackerBench v0.3 仅放行 3.3% 危险双用途提示。

9月21日周一
  1. xAI:News(网页)74

    xAI 发布 Grok 4.7,主打编码与知识工作

    xAI 发布 Grok 4.7,定位为其最强编码与知识工作模型,定价 $2/百万输入 token、$6/百万输出 token,与 Grok 4.6 同价同速,另有速度和价格加倍的快速变体。

    推荐理由:官方给出了完整定价、速度和多项基准对比,读者可以据此把 Grok 4.7 放进现有模型选型里横向比较。

9月20日周日
9月19日周六
  1. xAI:News(网页)64

    xAI 发布 Grok Voice Transcribe 2.0 语音转写模型,准确率较 1.0 翻倍

    xAI 发布 Grok Voice Transcribe 2.0 语音转写模型,官方称在真实场景评测中准确率是 1.0 的两倍且价格不变。在 Artificial Analysis 公开榜单上,它在 32 个 streaming 模型中准确率排名第一;多语言能力提升最大,短短语集合上词错误率从 20.6% 降至 6.8%。

    推荐理由:原文给出公开榜单排名、内部错误率数字和与 1.0 同价的定价,读者可据此评估迁移成本与实际收益。

9月17日周四
  1. xAI:News(网页)61

    Grok Build 推出记忆功能,可跨会话保留项目约定与决策

    xAI 宣布 Grok Build 上线记忆功能,会在每轮对话结束后于后台记录项目约定、决策及事实,供后续会话读取。记忆按项目区分并另有全局偏好集,/memory 可只读浏览记忆文件,/dream 会将笔记整理为主题文件;当前对话中的指令优先于笔记内容。

    推荐理由:原文说明了记忆的记录、整理和优先级机制,读者可以据此评估它对长期项目编码工作流的影响。

9月14日周一
9月13日周日
  1. IT之家(RSS)56

    AI 高管呼吁放缓前沿模型研发,分析称芯片股短期承压长期影响有限

    据彭博社报道,Anthropic CEO 达里奥·阿莫代伊呼吁全行业放缓最前沿模型研发,将增设独立第三方评估等安全防护措施,OpenAI CEO 奥尔特曼与马斯克表态支持。市场观察人士认为半导体及 AI 相关股票短期或遭抛售,但因芯片、能源与算力需求供不应求,长期影响大概率有限;纳斯达克 100 指数较 6 月高点已跌超 4%,美国芯片指数下跌 14%。

9月12日周六
  1. 🚨 AI News | TestingCatalog56

    Musk 回复称 Grok 4.7 还需几天继续打磨,可能因为在 RL 中对回复长度惩罚过重,模型仍会过早放弃其实能完成的困难任务,检查自身工作也不够严谨。TestingCatalog 补充称 9 月下半月 Meta 和 OpenAI 将在 OpenAI DevDay、Meta Connect 等活动前后发布模型,Grok 4.7 面临较高竞争门槛。

    引用Elon Musk@elonmusk

    @farzyness Grok 4.7 needs a few more days to cook. We might have penalized response length too much (or something) in RL, as it still gives up on hard tasks (that it can do!) too early and isn’t yet sufficiently rigorous in checking its work.

9月11日周五
  1. Elon Musk66

    Elon Musk 转发 Grok Bot 对 SpaceX CFO Bret Johnsen 在 Goldman Sachs Communacopia 演讲的摘要。

    引用DogeDesigner@cb_doge

    Grok Bot Summary of SpaceX CFO Bret Johnsen at Goldman Sachs Communacopia today. Vertical integration Vertical integration is the company’s core operating model, not a side strategy. - Rockets: own metal → engines → avionics → software - Starlink: own launch, satellites, and the end customer - AI: build facilities and power themselves, run their own models, sell to consumer and enterprise, and soon orbital compute Starship and launch Starship is the foundation for every other business. - Flight 13: big learning flight. Delivered demo V3 payloads, relit a Raptor, and got a soft, precise second-stage splashdown. Recovery team towed the stage back so engineers could study the heat shield. - Those learnings feed straight into Flight 14 and beyond. - Flight 14 (later this month): first revenue-generating Starship flight, flying production V3 Starlink satellites. - Later this year: aim to recover both first and second stages. Orbital compute Most of the AI industry agrees orbital compute is the future. Almost everyone else thinks it’s many years away. SpaceX disagrees because they control the stack. - Target: first orbital compute satellites next year - Scale: big compute in space into 2028 - Hardware approach: same V3 bus as Starlink, swap the payload, add larger solar arrays Why orbital can beat terrestrial on cost The crossover is about Starship reusability. - Falcon 9: first-stage reuse since Dec 2015; 500+ booster reflights - Starship: first stage already recovered/reflown; second-stage recovery progressing - Goal: reflight of both stages as soon as next year, which drops deployment cost sharply Terrestrial compute is getting more expensive (power, cooling, buildings, real estate). Orbital rides the opposite curve: cheaper rockets + better/cheaper satellites + scale. Johnsen said cost parity could come as soon as next year. Terrestrial compute and the $100B ARR goal - End of this year: on track for ~$100B ARR (annualizing the December number) - New update: another hosting deal closed earlier this month → about $1.1B/month starting Dec 1 → roughly +$13B ARR - Capacity: end this year well over 2 GW; next year 5–10 GW deployed - Confidence comes from line of sight to power, facilities, and permitting, plus being NVIDIA-exclusive for allocation - They stand compute up fast for themselves and for industry partners, which strengthens the NVIDIA relationship How they monetize compute Most hosting deals are short: ~90 days with a 90-day out (~6-month commits), including the newest deal. Why keep them short? - High conviction in their own products (Grok, Grok Bot, Cursor team after closing that deal) - Don’t want to lock forever capacity they may need internally - Internal bar: don’t let internal monetization fall below external hosting Earnings framing for next year: roughly $30–$50 per watt monetization range; they said they’re at the high end. Hosting customers appear to monetize even higher, which is why demand stays strong. Payback is under one year on new compute capex, so residual GPU value and financing options look attractive. “Not all CapEx is the same” — GPUs with <1-year payback are different from a launch tower built for decades. AI products and M&A Historically SpaceX was almost all organic growth. This year they did M&A because the AI product cycle rewards speed to frontier. - Closed Cursor deal weeks ago; product cycles already accelerating (called out Grok Bot) - Grok 4.6 improved on 4.5; 4.7 coming soon - Pitch: best infrastructure + competitive model + lower token cost = best position for customers - Market mood shift: months ago people bought the infra story but doubted the products; ~90 days later that skepticism is fading Starlink broadband Started as “better than nothing” (~2020–21). Now enterprise-grade with strong uptime/SLAs. - Resiliency pitch: boards will ask why Starlink wasn’t in the network if you go down - Mobility: aircraft backlog is large and production is ramping; cruise ships, yachts, trains too - Awareness, especially outside the US, is still a growth unlock - Longer-term: physical AI (robots, cars, aircraft) will need always-on connectivity terrestrial networks can’t fully cover Mobile / direct-to-cell Not a distraction. Same V3 bus, different payload. - Fly direct-to-device satellites through next year - Target service turn-on: first half of 2028 - V1 today (e.g. T-Mobile / T-SAT): text / light voice, great for emergencies and dead zones - Next gen: full 5G-quality from space - US: mid-band spectrum from EchoStar, FCC path for space + terrestrial - Go-to-market: flexible — own terrestrial build, or partner with carriers - International: same regulator-by-regulator playbook as broadband (Starlink now in 170+ countries) Near-term priorities: 1. Starship (enables everything else) 2. Terrestrial compute (funds growth and teaches them how to do orbital) Bottom line in one line Own the full stack, make Starship reusable at scale, use terrestrial AI compute as a cash engine now, and use the same satellite bus + Starship cadence to win broadband, mobile, and orbital AI.

    推荐理由:原文以 Grok Bot 摘要形式整理 SpaceX CFO 演讲要点,覆盖 Starship 复用、地面算力营收和轨道计算时间表等关键信息。

9月7日周一
9月6日周日
  1. 🚨 AI News | TestingCatalog52

    Grok Imagine Video 1.5 agent 现已可用,由最新的 Image 2.0 图像模型驱动,官方称质量更高、叙事更强,并擅长以更好的连续性衔接多个镜头。TestingCatalog 用 "Make me an ad banner for @testingcatalog" 实测后认为结果相当不错。

    引用Grok@grok

    Grok Imagine Video 1.5 agent is now available. Powered by our newest Image 2.0 model, it delivers higher quality, better storytelling from a smarter agent and excels at connecting multiple shots together with greater continuity.

9月5日周六
  1. lauren43

    Grok Bot 上线 Bot 模板市场,用户可直接添加创作者制作的模板,市场目前有 68 个公开 Bot、42 位创作者、10 个分类。官方分享的内部采购专员 Haggle Bot 可谈判供应商合同、找出闲置 SaaS 席位并比价经常性采购,上线一周为其团队节省超过 10 万美元。

    引用Grok Bot@bot

    You can now add Bot templates from our marketplace. We’re sharing Haggle Bot, our in-house procurement specialist. It negotiates vendor contracts, finds unused SaaS seats, and price-checks recurring purchases. One week in, it's saved us over $100K.

9月4日周五