What we measured, and what we still get wrong
Every figure on this page comes from a regression report generated by the test suite. Where we have not measured something, we say nothing rather than estimate.
Verification
Three tiers, three different failure modes
Each tier catches something the others cannot. Reference comparison catches a wrong rule. The determinism lock catches a rule change that moved an unrelated date. Audit replay catches a deployment that diverged from what was audited.
Tier A
42 / 42
Reference comparison
vs DrikPanchang
7 festivals × 6 years, spanning 1950–2100
Diwali, Janmashtami, Maha Shivaratri, Ganesh Chaturthi, Vijayadashami, Holi and Makar Sankranti are compared against DrikPanchang across a 150-year span. This catches rules that happen to be right for the current decade but wrong historically.
Tier B
Locked
Determinism lock
vs Itself
Every-day festival hash across 8 years
A hash of every festival on every day for eight years is stored in tier_b_lock.json. Re-running the engine must reproduce it exactly. This is the guard against silent drift: change a rule for one festival and accidentally move another, and the lock breaks immediately.
Tier C
26 / 26
Audit replay
vs Production data
22 days from the 2026–27 accuracy audit
Assertions taken from a manual accuracy audit are replayed against production. Where Tier A proves agreement with an external reference, Tier C proves the deployed system still behaves as audited.
Regression testing
152 tests, 91% coverage
The suite runs on every build. The coverage gate is 90%; below that, the build does not ship.
| Suite | Tests | Scope |
|---|---|---|
| festival_engine_v2/tests/ | 115 | Evaluators, kala, intervals, tie-break, series, validator, engine, db adapter, materialisation |
| tests/test_endpoint_tiers.py | 16 | Plan gating matrix, including Enterprise backward compatibility |
| tests/test_webhooks.py | 21 | HMAC signing, backoff, delivery outcomes, idempotency, producer filtering |
| Total | 152 | Coverage 91% on festival_engine_v2 (gate 90%) |
Why our coverage number went down
Performance
Measured on production, not a laptop
Ten samples against the live API. Engine-level figures are measured inside the festival engine, excluding network and serialisation.
End-to-end, production
| Endpoint | Median | p95 | Target |
|---|---|---|---|
| /v1/panchang | 138 ms | 154 ms | < 2000 ms |
| /v1/festivals | 148 ms | 169 ms | < 2000 ms |
| /v1/festivals/explain | 160 ms | 183 ms | < 2000 ms |
Engine-level
| Metric | Measured | Target |
|---|---|---|
| Festival lookup p95 | 2.4 ms | < 50 ms |
| Full year sweep | 0.8 s | < 5 s |
On the numbers you may have seen elsewhere
Coverage
What is verified, and how far
Festival derivation
location 1
Historical span
1950–2100
Determinism
8 years, every day
Ayanamsa
Lahiri
Corrections shipped
Defects we found and fixed
These were real errors in production before Festival Engine V2. We list them because a platform that claims never to have been wrong is not one you should trust.
| Festival | Was | Now |
|---|---|---|
| Diwali / Lakshmi Puja | Missing entirely | Amavasya prevailing at Pradosha |
| Krishna Janmashtami | Missing entirely | Nishita-vyapini Ashtami (Smarta default) |
| Ganesh Chaturthi | One day late | Madhyahna-vyapini |
| Vijayadashami | One day late | Aparahna-vyapini |
| Maha Shivaratri | One day late | Nishita-vyapini |
| Holi | Wrong when Purnima ended before sunset | Holika Dahan + 1 day, as a series |
| Abhijit Muhurat | Suppressed by Rahu/Yamaganda overlap | Always returned; overlaps listed separately |
Known limitations
What is still wrong, in our own words
This list is the actual contents of our internal limitations document. We would rather you find a limitation here than discover it in production.
amanta_month_key is mislabeled in source data
The column is off by one month for Krishna-paksha days. V2 rules therefore use the verified-correct purnimanta_month_key exclusively.
Panchang data coverage is uneven across locations
Locations 2–9 have no supplementary tables and fall back to live Swiss Ephemeris computation. Locations 10–15 are roughly 65% imported. Hyderabad has 51 of 701 years cached.
Rule packs exist but only north_india is populated
The pack mechanism is implemented and tested. Regional variants — Gujarati, Tamil, ISKCON, Nepal — are not yet authored.
Materialisation covers location 1 only
The nightly job materialises a 430-day window for location 1. Other locations compute on demand.
Webhooks are one-way and at-least-once
There are no inbound webhooks. Delivery is at-least-once, so subscribers must deduplicate on X-TathaAstu-Event-Id. Retries stop after 5 attempts; endpoints auto-disable after 20 consecutive failures.
Philosophy
How we handle being wrong
Rules are never edited in place
A correction inserts a new rule version pointing at the row it supersedes. The old row is marked DEPRECATED, not deleted. Any response we served in the past can still be reproduced from its recorded rule id and version.
Every correction is traceable to a reason
Traces record the rule, its version, the evaluator, the overlap in minutes, and the shastric authority. When a date changes, you can see precisely which rule changed and why.
Limitations are published, not buried
The list above is the actual contents of our internal KNOWN-LIMITATIONS document. We would rather you find a limitation here than discover it in production.
A number we cannot measure is not a claim we make
Every figure on this page comes from a regression report generated by the test suite. Where we have not measured something, we say nothing rather than estimate.
Report a date you believe is wrong
Send the request you made and the value you expected. Accuracy reports are triaged ahead of everything else. If you are right, the fix ships as a new rule version and appears in the changelog.