A benchmark of 27 AI models by Paulo Teixeira graded 17,850 answers and found prompt format barely moves accuracy — but ...