LogiSymb @ IJCAI-ECAI 2026
From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving
Dipankar Sarkar
Abstract
When an LLM-generated constraint program is unsatisfiable, returning a minimal unsatisfiable core instead of an error message localises the fault and cuts fabricated solutions from 79% to 7%.
Making language models solve constraint problems reliably often means having them translate the problem into a formal specification and delegating the search to a sound solver. But the translation is itself a language-model task, and an unfaithful translation makes the solver faithfully solve the wrong problem. Existing pipelines repair only translations that crash, returning the solver’s error message and falling silent when the program runs but is wrong. We replace the error message with a proof: when the generated program is unsatisfiable, we extract a minimal unsatisfiable core over the model’s own constraints and hand it back the exact set that cannot hold together, a leakage-free signal that localizes the fault. On a new benchmark of 77 problems with an exact oracle, translation to Answer Set Programming is faithful on six of seven domains and fails only on aggregate coverage scheduling, which concentrates the translation tax in one diagnosable pattern. A minimal core, rather than a bare error, is what stops a weaker model from fabricating solutions to infeasible problems, cutting fabrication from 79% to 7%. A strong chain-of-thought baseline meanwhile matches the symbolic route on accuracy, so the route’s value is not accuracy but certificates and its refusal to fabricate.
arXiv comments: 7 pages, 2 figures. Accepted at the IJCAI-ECAI 2026 Workshop on Logic and Symbolic Reasoning (LogiSymb), poster
Frequently Asked Questions
What is minimal-core-guided repair?
When the program a language model generates is unsatisfiable, the method extracts a minimal unsatisfiable core over the model's own constraints and returns that exact set to the model, localising the fault instead of returning a bare error.
What did it change?
On a benchmark of 77 problems with an exact oracle, it cut a weaker model's fabrication of solutions to infeasible problems from 79% to 7%. A strong chain-of-thought baseline matched the symbolic route on accuracy, so the route's value is certificates and refusal to fabricate rather than accuracy.