At the end of brev paper there is a section on the LLM usage.<p>Now, this paper will undergo peer review. And thus, when published, we will trust the result as well as any other published result in mathematics. However, while humans are not infallible, LLMs have a tendency to spit out confident stuff which looks correct. Thus I worry when it is used this way to produce proofs. Peer review is not perfect, and may not be tuned to catch LLM’s style of errors.<p>There is a solution to this, which is to formalise the result and get it machine verified. And with LLMs, I daresay this is going to be best practice moving forward.
> Peer review is not perfect, and may not be tuned to catch LLM’s style of errors<p>This summarizes I think a lot of the challenges with validating LLM output. We hear “humans make mistakes too”, but I would agree with you that our human detection of human-made mistakes and LLM-made mistakes is unlikely to have the same coverage.