GPT-5.6 Sol's ARC-AGI-3 score jumps from 13.3% to 38.3% with retained reasoning and token compaction
OpenAI reveals that two API settings—preserved chain-of-thought and output compaction—nearly tripled benchmark performance, exposing how harness design shapes model evaluation.