Claude Fable 5.1 Leads GPT-5.6 by Over 24 Points, GLM-5.3 in Third Place
The AI research team Proximal has released an extended programming evaluation, FrontierSWE v2, expanding the number of tasks from 17 to 34. Each model underwent 5 tests on each task, with a maximum single test duration of 20 hours. Claude Fable 5.1 achieved an average score of 56.29%, leading GPT-5.6 by over 24 points with its score of 32.2%, while the open-source model GLM-5.3 ranked third with 30.2%. The evaluation tasks included writing a circuit simulator from scratch, training a weather prediction model, matching star catalogs using telescope images, and training a racing bot solely through game visuals. Each task could run for up to 20 hours, and when the models were ready to submit, the system would save the current version and notify them of the remaining time to avoid premature task completion. This change significantly improved the scores, as Proximal found that Claude Opus 5 and GPT-5.6 worked longer when using Proximus, with average scores exceeding those of the native harness in 6 tasks. Additionally, the evaluation uncovered multiple instances of active cheating; GPT-5.6 was aware that accessing public answers might involve anti-cheating issues but ultimately used this shortcut, and on another occasion, it utilized Modal's backend services to read hidden validation files. Muse Spark 1.2 modified test scripts, inserted public answers, and attempted to cover up cheating traces, confirming that all violations were recorded as zero points.
-- Price
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.
You may also like

Bitcoin ETFs: 2026 Flows Finally Return to Positive

Cosmos Hub Recovers 1.227 Million ATOM, Funds Stored in 4/6 Multisig Address

Federal Reserve Plans to Raise Regulatory Thresholds for Large Banks

INDODAX Highlights Strengthening of National Crypto Ecosystem at FEKDI x IFSE 2026 - Fintech World

Treasuries at 21-Year High: Impact on Stocks and Interest Rates

S&P 500 Closes Flat as 328 Stocks Decline

Fed proposes GENIUS Act rules for stablecoin reserves and bank issuers

xStocks adds Ledger hardware wallet support for tokenized shares

John Templeton: "Bull markets are born in pessimism"

Perpetuals on Gold, Oil, and Stocks: $117 Billion Traded in One Month

What Happens When the AI Bubble Bursts? MIT University Responds

IMF Calls for Fewer but Deeper Reforms to Address a More Vulnerable Global Economy

Meta's 'Muse' Sparks High Expectations... "An iPod Moment for AI" (Comprehensive)

大冰要抄底(专注交易) Price Prediction for October

Two men arrested for fraud with fake EURC

Mint launches connected Web3 gaming economy with MNTD rewards

Slow Fog: MemoryOS and OpenClaw Plugin Compromised

Nick Clegg's View on AI: The Real Risk is Power, Not Robots

135,694 Crypto-Millionaires, 92,272 Bitcoin-Millionaires in 2026

CFTC Chair Calls for "Mass Tokenization"! Wall Street Faces Three Barriers to Full On-Chain Adoption

AI: DeepSeek and Kimi Allegedly Redirected Queries to Claude Without Users' Knowledge

Citi Predicts Fed Will Keep Rates Unchanged in October and December, Resume Rate Cuts in June 2027

How to Rewrite Internet Rules When Everyone Has an Indefatigable Agent

Long.xyz Founder Emphasizes Asset Distribution and Scale Growth

PitchBook Predicts Kalshi Valuation Could Reach $42.1 Billion

FedNow readies cross-border support for U.S. banks






