An expert TLA+ practitioner examines AI-generated formal specifications and finds a critical flaw: LLMs consistently produce only 'obvious' properties that are tautologically true or trivially satisfied, rather than the subtle invariants, liveness properties, and concurrency-related checks that make formal methods actually valuable. Using real GitHub examples of vibecoded TLA+ and Alloy specs, the author shows the specs don't compile and their assertions verify nothing meaningful. The post questions whether LLMs truly lower the skill barrier for formal methods, noting that expert specifiers can coax better results from LLMs — suggesting you may still need to know formal methods to use LLMs effectively for formal methods.

7m read timeFrom buttondown.com
Post cover image
Table of contents
No newsletter next weekLooking at a projectIs this a user error?Logic for Programmers Giveaway
7 Impressions