Birth-Time Rectification: A Blinded 23-Subject Backtest of a Vedic Dasha–House-Lord Engine
Rectification is the claim that an unknown or unreliable birth time can be recovered from a person's dated life events. It is also falsifiable: take a person whose birth time is independently known from a certificate, give the engine only their birth date, birthplace and eight dated events, and see whether it comes back near the known time. This study does exactly that — twenty-three public figures with certificate-grade birth records, put through the real production pipeline, scored against a pass rule fixed in the protocol before any run. The result: one out-of-sample pass in twenty-two, exactly what chance predicts, with the only near-hits explained by the engine's own design properties. The service will not open until a fresh cohort says otherwise, and this report is written so that practising astrologers can tell us what we are doing wrong.
Abstract
We backtested a Vedic birth-time rectification engine against twenty-three Rodden-AA subjects (certificate-grade birth records), eight or more sharply dated life events each, and a held-out recorded birth time. The engine searched the full birth day at one-minute resolution (1,440 candidates) using only birth date, place and event dates — the recorded time was never supplied — and scored each candidate with a classical Vimśottarī daśā + house-lord + karaka + Pranapada model (Pushya Paksha ayanamsa, whole-sign houses, mean lunar nodes). Pass rule (fixed in the protocol before any run): engine winner within ±30 minutes of the recorded time.
Finding: two of twenty-three subjects fall inside the ±30-minute band — Steve Jobs (Δ 9 min), who was the one subject used to tune the scorer and is therefore excluded from the out-of-sample count, and Albert II of Monaco (Δ 4 min), whose winner zone the engine's own diagnostics show to be structurally identical to the false "H1" farms — flagged, not certified. The out-of-sample count is therefore 1/22 (4.5%) against a no-skill null of 4.24% (p = 0.61 for one or more; the recorded time's mean rank, 788 of 1,440, is worse than the null expectation of 720). An independent timezone audit of every payload caught a sign error in two EDT births (Trump, Bush) that had produced a spurious third "pass"; corrected, that pass is a Δ 464-minute miss. The expert-recommended v2.0.0 Varga matrix (D-1 lordship corroborated in D-9/D-10/D-7) was implemented and tested both in-sample (it attacks the H1 misses but removes the only near-hits, adding no pass) and — the decisive test — out-of-sample on ten fresh Rodden-AA subjects never used in any tuning: every variant scores 0/10, identical to the baseline. The engine shows no above-chance rectification signal on out-of-sample data. It stays at v1.5.0, and the rectification service remains closed until a fresh cohort meets the protocol-fixed accuracy gate.
The claim and the test
Rectification is the foundation claim of our planned service: that a birth time can be recovered when the recorded time is unknown, unreliable, or — as recorded birth practice guarantees — only approximately right. The recorded time is rarely the whole story: it can capture head-out rather than full delivery, a scheduled caesarean rather than natural first breath, clock rounding, or a handwritten register. Rectification exists precisely to find the time the events point to. That makes the claim testable in the strongest sense: if the engine cannot find a known time from the same kind of evidence a customer would supply, it is not doing what it advertises. Twenty-two out-of-sample subjects is enough to test whether the engine recovers known times at all.
Method (protocol fixed before any run)
Subjects and blinding
Twenty-three public figures with Rodden-AA birth records: the original thirteen (Jobs, Ali, Elizabeth II, Trump, Jordan, Obama, Bush, William, Victoria, Alice, Turner, Albert II, Stéphanie — runs 1–13) plus a same-day expansion of ten fresh subjects never used in any tuning (Diana, Charles III, Elvis Presley, Marilyn Monroe, JFK, Churchill, Einstein, Reagan, George VI, Biden — runs 15–24). Each subject contributes eight or more sharply dated, independently documented events across several life areas (career, marriage, childbirth, accident). Blinding: the known time lives in the subjects file and is never sent to the engine — the API call carries only birth date, place (coordinates + timezone), and the events as {type, date, text}. Every birth record was cross-checked against Astro-Databank; the collaborating model's first draft contained five wrong birth times/dates, all corrected against AA entries before any chart was computed.
Engine and parameters
For each of the 1,440 candidate minutes the engine computes the Vimśottarī
Mahādaśā and Antardaśā lords active at each event date and awards points when those lords
hold a classical claim on the event: ruling the event's house from the candidate's
lagna (marriage → 7th, childbirth → 5th, career → 10th, accident → 8th); being the
event's natural significator (karaka); or sitting in the event's house natally. Reproducible
parameters: Pushya Paksha ayanamsa
(get_pushya_paksha_ayanamsha), whole-sign
houses, mean Rahu/Ketu (the true-node
option is not used), one-minute search grid, deterministic scoring — no LLM in the ranking
path. Three scoring refinements came out of a 120-run step ablation (measured, not
guessed): per-type normalisation, MD==AD dedupe, and the Pranapada micro-layer — the
largest single contributor to the in-sample score during tuning; on the out-of-sample
cohort it produces no recovery (median |Δ| ≈ 7 h), so it is a tuning-sensitive term, not a
general accuracy driver. The daśā skeleton is nearly invariant across the day (the Moon's
nakshatra fixes the daśā sequence), so almost all discrimination comes from the house-lord
frame — from which lagna each candidate produces. That single fact is the root of the
failure mode analysed below.
Tuning disclosure
Steve Jobs (run 1) was the subject used to tune the scorer during the 120-run step ablation. His Δ 9-minute result is therefore in-sample and carries no evidential value; he is excluded from every out-of-sample statistic in this article. The ten-subject expansion was assembled after tuning and is fully out-of-sample.
Pass rule and statistics
Pass = engine winner within ±30 minutes of the recorded time; under the no-skill null the per-subject pass probability is 61/1440 = 4.24% (binomial across subjects). A second, distribution-free metric is the rank of the recorded time among the 1,440 candidates (null expectation 720.5). The recorded time is an anchor, not ground truth — the protocol therefore judges the cohort distribution (median |Δ|, % within bands, systematic direction bias, explainable outliers), never curve-fits Δ → 0 on a single chart.
Results — the 23-subject cohort
| Subject | Recorded | Engine | Δ | Verdict |
|---|---|---|---|---|
| Albert II † | 10:50 | 10:54 | 4 min | ✅ PASS — H1-flagged (rank 849/1440, Δ-coincidence on an H1-shaped farm) |
| Steve Jobs ‡ | 19:15 | 19:24 | 9 min | ✅ PASS — tuning subject, excluded from out-of-sample counts |
| Michael Jordan | 13:40 | 15:38 | 1h 58m | ❌ MISS |
| Marilyn Monroe | 09:30 | 11:36 | 2h 06m | ❌ MISS |
| Stéphanie | 18:25 | 21:03 | 2h 38m | ❌ MISS |
| Charles III | 21:14 | 00:29 | 3h 15m | ❌ MISS |
| George W. Bush | 07:26 | 11:54 | 3h 32m | ❌ MISS |
| Muhammad Ali | 18:35 | 14:02 | 4h 33m | ❌ MISS |
| Elvis Presley | 04:35 | 22:49 | 5h 46m | ❌ MISS |
| Barack Obama | 19:24 | 13:28 | 5h 56m | ❌ MISS |
| Ronald Reagan | 04:16 | 21:46 | 6h 30m | ❌ MISS |
| Ted Turner | 08:50 | 15:39 | 6h 49m | ❌ MISS |
| Joe Biden | 08:30 | 01:35 | 6h 55m | ❌ MISS |
| George VI | 03:05 | 20:07 | 6h 58m | ❌ MISS |
| John F. Kennedy | 15:00 | 22:15 | 7h 15m | ❌ MISS |
| Prince William | 21:03 | 04:22 | 7h 19m | ❌ MISS |
| Donald Trump § | 10:54 | 03:10 | 7h 44m | ❌ MISS — tz-audit retraction (was a 16-min "pass") |
| Diana | 19:45 | 04:16 | 8h 31m | ❌ MISS |
| Queen Victoria | 04:15 | 12:59 | 8h 44m | ❌ MISS |
| Albert Einstein | 11:30 | 01:12 | 10h 18m | ❌ MISS |
| Princess Alice | 04:05 | 17:24 | 10h 41m | ❌ MISS |
| Winston Churchill | 01:30 | 12:14 | 10h 44m | ❌ MISS |
| Elizabeth II * | 02:40 | 14:33 | 11h 53m | ❌ MISS (weak anchor — caesarean) |
† H1-flagged · ‡ tuning subject (excluded) · § tz-audit retraction · * weak anchor. Δ times use the shortest path around the clock (e.g. 21:14 → 00:29 = 195 min). Runs 1–13: original cohort (2026-09-01). Runs 15–24: expansion (2026-09-05), ten fresh Rodden-AA subjects — 0/10 MISS.
Is the result better than chance?
| Test | Data | p |
|---|---|---|
| Exact binomial, ≥2 passes in 23 (protocol-fixed rule, includes tuning subject) | all subjects | 0.255 (ns) |
| Exact binomial, ≥1 pass in 22 | out-of-sample (Jobs excluded — tuning data) | 0.614 (ns) |
| Exact binomial, ≥2 passes in 22 | out-of-sample | 0.239 (ns) |
| Sum-of-ranks of the recorded time (Monte-Carlo) | out-of-sample 22 | 0.776 (ns) |
| Top-10% count of the recorded time | 2 vs 2.3 expected | ≈0.69 (ns) |
Every test is null. The out-of-sample pass rate (1/22 = 4.5%) sits on top of the chance expectation (4.24%); the recorded time's mean rank (788) is worse than the null (720); the median displacement is 415 minutes. The two within-band results do not rescue the cohort: Jobs is tuning data, and Albert II's winner — examined below — is the same object as the false positives, 4 minutes away by coincidence. This is an existence failure, not merely a reliability one: on 22 out-of-sample certificate-grade subjects the engine recovers a known birth time no better than a random minute of the day.
Failure analysis — H1 is a design property, not a bias
Across the misses the same signature appears: the engine prefers a candidate lagna where the event-era daśā lords happen to rule the event houses, hours away from the recorded time. This is not a bias that crept into the implementation; it is what the scoring function does by construction. The daśā skeleton is fixed for the day, so the only time-varying degree of freedom is the lagna; the scorer therefore finds the lagna that maximises the house-lordship term, and because a D-1 lagna changes only every ~2 hours, a career cluster inside a long Mahādaśā gives the engine a large target area in which the era lord rules the 10th. The equation is under-constrained — the lordship credit is not corroborated by any faster-changing frame. Examples:
- Victoria (the purest demonstration): the same Mercury-MD era scores 3.0 + 4.2 for both Jubilees from the false winner's lagna and 0.0 + 0.0 from the recorded one — no timeline shift, purely the house-rulership lottery.
- William: Jupiter-MD covers 2006–2018 at the false winner and rules the 7th/10th, so marriages score 6.2 and careers 5.0–6.2; at the recorded time those events fall under Saturn-MD, which rules neither house → 0.0 on both marriages and all three childbirths.
- Obama and Bush: from the false winners' lagnas the era lords rule the 5th (childbirths) and 8th (9/11) respectively; at the recorded times all events are karaka-only.
- Boundary-mining: many false winners sit at sign boundaries (Obama Scorpio 0.2°, Bush Gemini 0.1°), where a one-minute change discontinuously reassigns every house lord — the coincidence score spikes at exactly the minutes that are least stable.
Albert II is the same object. The engine's per-event diagnostics show his winner zone is an H1-shaped farm — one Jupiter-era lord ruling the 10th and 7th and carrying childbirth karaka across 2003–2014 — structurally indistinguishable from William's false farm; only the distance to the certificate differs (Δ 4 min vs Δ 439 min). His recorded time ranks 849th of 1,440 by score, so the "pass" is a Δ-coincidence on an under-constrained score surface, not a certified signal. He is flagged accordingly in the cohort table and not counted as evidence.
Attempted fixes — measured and rejected
Every candidate fix below was implemented and run against the cohort before acceptance or rejection (deterministic engine, zero LLM calls in the scoring path).
The expert v2.0.0 roadmap (2026-09-04)
An external expert review diagnosed H1 as an under-constrained equation and prescribed the classical remedy: the Varga matrix — score each event in D-1 and the event's divisional chart (marriage → D-9, career → D-10, childbirth → D-7), whose lagnas change every ~12–13 minutes, so a house-lord coincidence that survives four hours in D-1 almost never survives the varga window — plus a sandhi penalty on lagnas within 0.5° of a sign boundary, and Tier-1 event weighting (marriage/childbirth ×1.5). All three were implemented as an 8-variant ablation over the 13-subject cohort.
| Variant (13 subjects, in-sample) | Mean rank | Passes |
|---|---|---|
| v1.5.0 baseline | 576 | 2/13 (Jobs ‡, Albert †) |
| v2-varga-gate (hard) | 531 | 0/13 — destroys both near-hits |
| v2-gate-half / v2-gate-mdonly | 521–541 | 1/13 (Jobs only) |
| v2-career-only | 562 | 2/13 (keeps both, adds none) |
| + sandhi penalty | 572 | 0/13 — breaks both boundary births |
| v2-varga+tier | 546 | 0/13 |
The decisive test: v2.0.0 out-of-sample (2026-09-05)
The in-sample ablation above is open to exactly the objection the reviewer raised: it judges a fix against a cohort that contains the tuning subject and an H1-flagged near-hit. The correct test is a cohort never used in any tuning. The ten expansion subjects were therefore run through all eight v2.0.0 variants with the identical harness:
| Variant (10 fresh subjects, out-of-sample) | Mean rank | Passes |
|---|---|---|
| v1.5.0 baseline | 985 | 0/10 |
| v2-varga-gate (hard) | 936 | 0/10 |
| v2-gate-half | 962 | 0/10 |
| v2-gate-mdonly | 952 | 0/10 |
| v2-career-only | 963 | 0/10 |
| v2-career+sandhi | 934 | 0/10 |
| v2-varga+sandhi | 934 | 0/10 |
| v2-varga+tier | 919 | 0/10 |
The Varga roadmap is refuted on data that cannot be accused of protecting Jobs or Albert. Every variant recovers zero recorded times on the ten fresh subjects — identical to the baseline. The varga gate does move the H1 misses' mean rank (the reviewer's direction is right: Victoria 369→253, Alice 476→324 in-sample), but movement toward the top is not recovery, and no variant produces a single pass. The sandhi penalty is refuted outright: it penalises exactly the boundary births rectification exists to resolve. The varga-witness fraction of a winner carries zero separation (genuine Jobs = 0.0 while false Ali/Jordan/Turner = 1.0), which closes the discriminator design space. The D-1 dasha-house-lordship term cannot be separated into genuine-vs-coincidental credit by any scoring-side construction we can build.
Other measured attempts
| Attempt | Result |
|---|---|
| Per-event LLM signature interpretation | Rejected — widened the coincidence surface |
| Lagna-rotation debias | Refuted — ties on mean rank, no pass gained |
| Analytic drift-detrend prior | Wash — no rank improvement |
| Karaka up-weighting | No change — confirms H1 lives in house-lordship credit |
| Jupiter/Saturn double transit (scored bonus) | Rejected — mean rank 546→657, removes the near-hits |
| 120-run step ablation (24 compositions × 5 subjects) | v1.5.0 composition best; Kunda gate removed as noise |
| Fewer-events sensitivity (2,340 subsets) | Albert II robust to thinning; Jobs pins from 7 events; misses structural — thinning never rescues them |
| Forced-type national events (9/11 as personal accident) | Bush improves 268→106 min without it — "never force an unnatural event type" logged as a product rule |
Data-quality audit: timezones and anchors
Every payload's timezone offset was re-verified against the engine's
convention (UTC = local − tz_offset, astro_engine.get_julian_day).
Two of the original thirteen payloads were wrong — Trump
and Bush, both EDT births (UTC−4), carried
tz_offset = +4.0 instead of −4.0, so their charts were computed
8 hours early: Trump's reported "11:10 winner, Δ 16-min PASS" was in fact the chart of
03:10 EDT (pre-dawn, Aries rising) compared across two mismatched clocks; corrected, it is a
Δ 464-minute miss. Bush was a miss in both frames (Δ 268 → 212 min). The historical zones
of the UK/European subjects were verified against DST tables rather than assumed: Elizabeth
II (21 Apr 1926 — British Summer Time in force, recorded 02:40 BST/UTC+1), Diana (1961,
BST), Charles III (Nov 1948, GMT), George VI (Dec 1895, GMT), Victoria (1819) and Churchill
(1874) — both pre-DST local-mean-time births; Biden was zoned as US War Time (EWT, UTC−4,
1942–45), the same trap class as the Trump error. The payloads are corrected in place and
all cohort numbers in this article use the corrected values.
Limitations
- Out-of-sample n = 22. Sufficient to show the engine performs at chance; insufficient to bound a small real effect tightly.
- Tuning disclosure. Jobs (run 1) tuned the scorer; his pass is in-sample and excluded. The ten expansion subjects are fully out-of-sample and all missed.
- Anchor error is real but cannot explain hours. Recorded ≠ destined (caesarean, rounding, recording practice). Elizabeth II is flagged as a weak anchor (lagna 0.29° from a boundary) — but systematic multi-hour displacement is an engine property (H1), not a data-quality artefact.
- Ayanamsa sensitivity. The engine uses Pushya Paksha ayanamsa (the production setting). Ayanamsa choice shifts the zodiac ~24°, and different ayanamsas place boundary-sensitive lagnas and varga divisions differently; no other ayanamsa was tested, and that is a stated boundary of these results.
- Narrow event vocabulary. Six event types (marriage, childbirth, career, accident, move-abroad, spiritual); everything else is mapped or excluded.
- No transits, no finer layers in the shipped engine. The double-transit and micro-pin layers were tested as scored terms and rejected; they remain untested as hard elimination gates.
- The LLM stage never scores. It interprets the engine's winner for the product report; it is deliberately kept out of the ranking.
Conclusion — go/no-go
On 22 out-of-sample certificate-grade subjects the engine recovers a known birth time at exactly the chance rate (1/22 = 4.5% vs 4.24%, p = 0.61), and its winner is not even concentrated near the truth (recorded-time mean rank 788 vs null 720; median |Δ| 415 min). The two within-band results do not constitute evidence: one is tuning data, one is an H1-shaped coincidence. The expert-prescribed Varga matrix — implemented and tested on a cohort that cannot be accused of protecting either — recovers zero times. The engine stays at v1.5.0, and the rectification service remains closed. This is not a claim that rectification is impossible or that the classical framework is wrong; it is a claim that this implementation shows no recoverable signal on fresh data, and that no scoring-side construction we have measured changes that. The next step is not another re-weighting of the D-1 lordship term but a layer independent of it — a lagna-hour prior built from the event distribution, or a micro-pin such as the Tattva/hour-ruler — validated against a fresh ≥20-subject cohort, never used in tuning, before it may touch the engine.
Study timeline
| Date (2026) | Step | Outcome |
|---|---|---|
| 08 | Protocol written (pass rule fixed before any run, blinding, anchor doctrine) | Design locked before any run |
| 09-01 | Runs 1–13 (engine v1.0.0 → v1.5.0, tuned on Jobs); 120-run step ablation; H1 named; sensitivity runs | 3 PASS / 10 MISS pre-audit; H1 reproduced in 11/11 misses |
| 09-02 | Cohort statistics + first external review (AI Studio) | "Weak signal, conclusive non-viability" |
| 09-04 | Expert review answers Q1–Q8 → v2.0.0 roadmap; roadmap implemented as 8-variant ablation + era-fit follow-up | All forms refuted in-sample; engine stays v1.5.0 |
| 09-05 | Timezone-sign audit; cohort expansion (10 fresh AA subjects); v2.0.0 re-run on the fresh 10 (out-of-sample) | Trump's "PASS" retracted (Δ 16→464); 0/10 expansion; v2.0.0 = 0/10 every variant — cohort at chance |
Reproduction
The complete study is published with this article: the protocol, committed to the repository before any run, the subjects file with birth-data sources and Rodden ratings, every run payload and result, per-subject diagnoses, the 120-run step-ablation output, the significance and audit notebooks, the v2.0.0 ablation notebooks (in-sample and out-of-sample), and the versioned engine changelog (v1.0.0–v1.6.0) so any result traces to the exact code that produced it. The engine used in every run is the exact engine a shop customer's rectification would run, minus the LLM verdict on misses, which are diagnosed deterministically.
This study is scientific research on the measurable claims of a planned service, published by the ATRAC Institute's Sidereal Telemetry Lab. It is published in full — including the negative result — because that is the standing rule of the research programme. The rectification service remains closed to customers until the protocol-fixed accuracy gate is met.