OpenAI removes text-chat limits for free ChatGPT users, debuts GPT-5.6 variants with reasoning controls
OpenAI rolls out unlimited text chats to free users and introduces GPT-5.6 Luna and Sol models with adjustable reasoning sliders, while reducing factual errors by 62–68% versus GPT-5.5-Instant.
GPT-5.6 Sol's ARC-AGI-3 score jumps from 13.3% to 38.3% with retained reasoning and token compaction
OpenAI reveals that two API settings—preserved chain-of-thought and output compaction—nearly tripled benchmark performance, exposing how harness design shapes model evaluation.
Thinking Machines Lab Launches Inkling, a 975B-Parameter Open-Weight Model Built by OpenAI Defectors
The startup founded by former OpenAI executives releases its flagship multimodal model, challenging the dominance of closed-source systems.
Thinking Machines Releases Inkling, a 1-Trillion-Parameter Multimodal Open Model
Inkling combines native image, audio, and text processing with a 1M-token context window and sparse MoE architecture for efficient multimodal reasoning.
OpenAI Launches GeneBench-Pro to Test AI Judgment in Computational Biology
OpenAI introduces GeneBench-Pro, a 129-problem benchmark measuring whether AI models can make higher-order scientific judgments in genomics and translational medicine.
Sherlock Holmes Board Game Becomes LLM-Agent Evaluation Framework
Researchers use a classic mystery game to test how well AI agents reason through multi-step deduction and suspect elimination.
Large Language Models Retain False Information Despite Explicit Warnings
Research shows LLMs incorporate contradictory statements into reasoning, even when explicitly told the claims are false.
Anthropic releases Claude Opus 4.8 with improved uncertainty flagging and effort controls
Claude Opus 4.8 flags uncertain reasoning 4x more often than its predecessor and introduces user-controlled effort levels and dynamic workflow agents.