The Reasoning Fragility Probe: Implementing Tests on the Limitations of Mathematical Reasoning in LLMs
There’s an assumption behind every benchmark score reported for LLMs that the model is doing something we’d recognise as reasoning. Not pattern-matching, not memorisation, not sophisticated autocomplete — actual reasoning. The kind where you understand the structure of a problem and apply it to a new instance, regardless of what the numbers look like. A…