Solved
932
65.6% coverage
Research in progress
Observatory is preparing and verifying real benchmark data. Public numbers will be released only after that verification is complete. Placeholder and test results stay hidden until then.
Logic capability
Deep analysis across 1,420 official benchmark instances, with correctness and runtime breakdowns.
Solved instances grouped by wall-clock runtime
Where each stable artifact stands
New capability and regressions exposed by the current run
| Benchmark | Logic | Result | Previous | Runtime | Features |
|---|---|---|---|---|---|
QF_UFNIA/scheduling/uf-mul-233.smt2 | QF_UFNIA | timeout | sat | 20.0 min | |
QF_UFNIA/locks/model-071.smt2 | QF_UFNIA | unsat | timeout | 22.10 s |