3 comments
>We study a frontier LLM in its purest inference regime - a single generation pass over a prompt containing the full training and test data, with no tools, no agentic scaffolding, and no fine-tuning<p><i>Sigh</i>. So this is somewhat interesting niche academic research but utterly irrelevant to real-world use cases.
Unsure if it's LLMs that fail at tabular data or its just that tree boosting are spectacular at that task.
Just have 2 LLMs debate whether tabs or spaces are the superior choice