← Back to the research library
Empirical Study

Personality and the Chart: A Blind Biography–Chart Matching Study Across Seven Language Models

September 2026By ATRAC Institute

If a person's character is written in the sky at the moment of birth, then a reader who knows nothing but that person's life story should be able to recognise their chart — and a reader who knows nothing but the chart should be able to recognise the life. That is the oldest and most testable claim in astrology, and modern language models finally make it possible to run blind at scale. This study runs the classic matching experiment in its most rigorous form yet — the protocol fixed in writing before any chart was computed, blinded at every step, and with the experiment's single greatest confound, the reader's own memory of famous people's charts, measured directly and removed by construction in a second round. The honest result: across six independent model families, an anonymised biography cannot be matched to its sidereal chart better than chance. Design, execution and analysis all took place within a single day (2026-09-04) — an exploratory pilot, not a formally pre-registered study, and it is presented as such.

Abstract

We tested whether sidereal chart factors carry a readable, biography-consistent personality signal by giving large language models a forced-choice matching task: ten anonymised biographies (≈300 words each, written blind to astrology) against ten sign-level chart abstracts (lagna · Moon · Sun · 10th house + occupants · exaltation/debilitation), computed with the production engine (Pushya Paksha ayanamsa, whole-sign houses) from independently verified Rodden-AA birth records. Endpoints were fixed in the written protocol before any run: primary = exact 1:1 hits (Montmort derangement null, ≥4/10 needed for p < 0.05), secondary = top-3 concentration and mean rank. Seven models from six independent families ran the identical task.

Finding: six independent model families (ChatGPT, Claude, DeepSeek, Grok, MiniMax — plus two failed Nemotron runs) score 0–2/10 exact, every p(A) ≥ 0.26: statistically indistinguishable from chance. Only the two Gemini-family models clear the bar (Flash 10/10, Pro 8/10, p < 0.001) — and direct memory probes showed why: both models carry the canonical celebrity charts in training memory (bio-side signature recall 5–8/10 at high confidence), while a memory-only simulation from their own recalled signatures resolves at most 5/10. Their scores are contaminated on the famous core and cannot certify a readable signal.

Round 3 removed the chart-content memory channel by construction: ten 2025 Nobel laureates whose charts could not have been published before the matchers' January-2025 knowledge cutoff, matched on date-only factors (Moon/Mercury/Venus/Mars signs, no lagna, no Sun), after a cleanliness gate showed both matchers can name the people (10/10, 9/10) but recall 0/10 chart factors at 0% confidence — the identity channel is open, the chart-content channel is not. Every Jan-2025-cutoff matcher scored 0/10 (pure chance) despite naming the subjects; the single above-chance run (10/10) is attributed to **Gemini Flash 3.8**, the one model whose March-2026 cutoff postdates the October-2025 Nobel announcements — so its training window contains either the post-announcement chart content about the new laureates or the publicised birth dates from which the date-only signs are derivable. Both are training-data channels, not chart-reading, and the 10/10 is not evidence. Conclusion: H1 is not supported — no reproducible above-chance biography↔chart matching signal was found in any round. The contribution is the method: a blinding-audited protocol, fixed in writing and committed to the repository before any run — an exploratory same-day pilot, not a formally pre-registered study — whose contamination channel is measured, named, and closed. Published in full, negative result included.

Why matching, and why now

Personality claims are the most falsifiable part of astrology. They make specific, observable predictions — a Moon in this sign, a Mars in that house, a lagna of a certain quality — about how a person behaves, decides, and presents. The cleanest experiment is matching, not prediction: hand a reader a set of anonymous biographies and an equal set of charts, shuffled, and ask for the 1:1 pairing. Chance is exactly 1 in N! — astronomically small for ten subjects — so even a few correct pairings are decisive if the task is truly blind.

The classical barrier was blinding. Human astrologers who read the famous biographies can rarely un-know the charts they have studied for decades, so matching studies leak through the expert's own memory. Language models remove the practical barrier but replace it with a subtler one: the same training data that makes a model a capable reader of chart symbolism also contains astrology websites publishing the sidereal charts of the most famous people on Earth. A model that has seen "Einstein — Gemini lagna, Scorpio Moon" in training is answering from memory, not reading. The study therefore treats training-data memory as the first-class confound it is: Round 1 measures the whole task honestly and flags the anomaly; Round 2 probes the memory channel directly; Round 3 rebuilds the experiment on subjects whose charts cannot be in any matcher's memory.

Design (protocol committed 2026-09-04; all runs same day)

  • Subjects. Ten public figures with Rodden-AA birth records (birth certificate or equivalent): Einstein, Hepburn, Jobs, Obama, Monroe, Hendrix, Cobain, Trump, Bell, Christie. The collaborating AI's first draft contained five wrong birth times/dates; all were corrected against Astro-Databank AA entries before any chart was computed.
  • Blind biographies. ≈300 words each, written by Gemini Flash from a strict archivist prompt: no names, no dates, no places, no gendered pronouns, no astrological vocabulary. Audited; labels A–J fixed; chart numbers shuffled separately (seed 20260904).
  • Chart abstracts. Sign-level only — lagna, Moon sign, Sun sign, 10th-house sign + occupants, exaltation/debilitation states. Nakshatras, degrees, vargas, dashas and outer planets withheld: the reader gets what a personality reader should need, nothing that could betray identity by computation.
  • Endpoints. A (primary): exact 1:1 hits, Montmort (derangement) null, ≥4/10 for p < 0.05. B (secondary): true bio in top 3, binomial p = 0.3. C (secondary): mean true rank vs 5.5 by permutation. Supportive metrics never override the primary.
  • Honesty rules. Every run is kept verbatim (raw responses archived). A run is invalid only if the model names the person behind a bio or provably holds the charts in training (cutoff evidence); invalid runs are recorded, never deleted, and analysed. Above-chance runs are never discarded — they are probed (§9.2), replicated, or set aside with the evidence that their model holds the subjects in training. A shuffled-chart control is specified in advance to adjudicate any above-chance score.

Round 1 — famous subjects, seven model families

The identical self-contained payload (bios + shuffled abstracts + instructions) was run on seven models from six independent vendors. Two early Nemotron free-tier runs are documented as invalid — the model recognised the people behind the anonymised bios and reasoned from its memory of their charts instead of the supplied abstracts, exactly the leakage the design anticipated; kept as evidence that models do know these ten charts from training. The rule is applied uniformly across vendors: recognition-based reasoning disqualifies a run whether the recognition is stated outright (Nemotron) or inferred from cutoff evidence (Flash 3.8 in Round 3).

Model family Exact (A) p(A) Top-3 (B) Mean rank (C)
Grok Fast0/101.0003/105.80
ChatGPT Think2/100.2645/105.10
Claude 5 High2/100.2644/105.00
DeepSeek Expert2/100.2644/104.40
MiniMax M3 (free)2/100.2647/103.70
Gemini Flash 3.810/10<0.00110/101.00
Gemini Pro 3.18/10<0.00110/101.20

Shuffled-chart control p for the null models: 0.264–0.267 (identical rankings scored against 10⁵ random truths — indistinguishable from chance). MiniMax's supportive metrics (top-3 7/10, p = 0.011; mean rank 3.70, p = 0.026) are each significant in isolation but the protocol forbids rescuing a null primary with supportive metrics, and no other model reproduced the concentration.

The distribution is the finding: six independent families converge on 0–2/10 — five of them at exactly 2/10 — while only the two Gemini-family models clear the bar. A genuinely readable signal would produce partial agreement across families; rank-1 agreement across families is 1–4/10, at chance. What the distribution looks like instead is a training-corpus effect: the Gemini pair is the one that can answer from memory. Round 2 was designed to decide exactly that.

Exact rank-1 matches per model in Round 1 (famous subjects) and Round 3 (chart-content-closed subjects), against the protocol-fixed ≥4/10 significance bar
Every Jan-2025-cutoff matcher in Rounds 1 and 3 sits at or below 2/10 — chance. The three above-bar scores are all Gemini runs whose training predates or postdates the subjects' chart content (hatched — memory artefacts, not evidence): Flash 10/10 and Pro 8/10 in Round 1 (canonical charts in training), Flash 3.8 10/10 in Round 3 (post-announcement chart content inside its March-2026 cutoff). Bottom pairs: Round 4's tropical A/B — MiniMax M3 1/10 sidereal vs 0/10 tropical, DeepSeek 0/10 sidereal vs 1/10 tropical — both families at chance in both zodiacs; no zodiac advantage (pooled 1/20 each).

Round 2 — measuring the reader's memory

Two memory probes were run on the Gemini pair, in both directions. The abstract side is the purest chart→person test: given only a chart abstract, name the person. If the models held a memorised chart→person lookup, this should be near-perfect. It is not: Gemini Flash names the right person 4/10, Gemini Pro 2/10 — and both answer confidently wrong on several (Flash: "Oliver Cromwell" for Hepburn; Pro: "Kurt Cobain" for Bell, "Narendra Modi" for Christie).

The bio side is the opposite direction: given the biography, name the person and recall their chart signature. Both identify all ten people trivially (the bios are de-anonymisable by design — that is the leak), and recall the canonical celebrity chart factors well: strong signature recall on 5/10 (Flash) and 8/10 (Pro) for exactly the subjects they also matched at rank 1 — Einstein, Obama, Monroe, Trump — citing their real training sources (BV Raman, KN Rao, astro-databank). A memory-only simulation — matching purely from each model's own recalled signatures — resolves at most 5/10 (Flash) and 4/10 (Pro), far below their actual 10/10 and 8/10.

Verdict: mixed, not clean. Pure memorisation is rejected (the abstract side fails, and the memory-only simulation falls half short — Flash additionally matched 5/5 on charts it explicitly said it holds no confident memory of, with factor-specific rationales). But the memorised core (Einstein/Obama/Monroe/Trump) is real and already clears the ≥4-hit bar by itself, and Pro's two matching errors are exactly its two memory gaps (Bell: "don't know"; Jobs: wrong lagna) — the fingerprint of memory-driven matching on the canonical subset. The Gemini scores are therefore disqualified as clean evidence on famous subjects, while the non-canonical residual stays consistent with genuine reading but uncertifiable in this design. The decision escalates to Round 3: subjects whose charts exist in no training corpus.

Round 3 — post-cutoff subjects (2025 Nobel laureates)

The leak is not fame — it is published astro content in training. So the fix is not "less famous people", it is "people whose charts were never published before the matcher's knowledge cutoff". The 2025 Nobel class (announced 6–13 October 2025) is ideal: any chart pages about them could not exist before the January-2025 knowledge cutoff of the matcher models. Two caveats are worth stating plainly. First, the laureates were already famous scientists, so their identities — and public birth dates — are in training; Round 3 closes the chart-content channel, not the identity channel, and the audit below keeps that honest. Second, because the factor set is date-only, a model that recognises a subject and knows the birth date could in principle derive the signs by ephemeris computation rather than by reading a chart — the reason the Sun (derivable from a public birthday) is excluded. All ten laureates have verified public birth dates but no published birth times, so the factor set is date-only: Moon, Mercury, Venus and Mars signs — fixed by birth date alone. Lagna is excluded (requires an unpublished birth time). Where a planet could sit in either of two signs depending on the hour of birth, both candidates were listed in the abstract with their crossing times held internally.

Cohort (2025 Nobel field) Factor set Ambiguity handled
Robson, Clarke, Mokyr, Howitt, Sakaguchi, Kitagawa, Devoret, Aghion, Ramsdell, YaghiMoon · Mercury · Venus · Mars signs6/10 fully date-fixed; 4 carry one hour-dependent factor (both signs listed)

Cleanliness gate (protocol §11.4), run first: both matchers were shown the blind bios and asked to name the person and recall any chart factors. Both matchers named nearly everyone (10/10 and 9/10) — expected, since the bios describe public careers — but recalled 0/10 chart factors at 0% confidence. The naming result is not a failure; it is the audit's point. The identity channel is demonstrably open, so the four verifiable Jan-2025 runs scoring 0/10 while naming the people is direct evidence that knowing who someone is does not, by itself, yield their date-only signs — and any above-chance run must be judged against models that provably recognise the subjects.

Matches: Gemini Pro 3.1 = 0/10 exact (p = 1.0); Gemini Flash 3.5 re-run at temperature 0 = 0/10 (p = 1.0); Gemini Flash 3.1 Lite = 0/10; Gemini Flash 3 Preview = 0/10 — four mutually independent Jan-2025-cutoff runs at pure chance, with zero agreement between them on which charts they approach. The fifth run, attributed to Gemini Flash 3.8 (session mix-up at run time), scored 10/10 — explained by its March-2026 cutoff, which places the subjects inside its training data (see below). The first four are the texture of noise; the fifth is the memory channel re-opening by accident, which is exactly why the gate's cutoff verification matters.

Round-3 run Exact (A) p(A) Top-3 (B) Verdict
Gemini Pro 3.10/101.0003/10null — chance
Gemini Flash 3.5 (re-run, temp 0)0/101.0002/10null — chance
Gemini Flash 3.1 Lite0/101.0004/10null — chance
Gemini Flash 3 Preview0/101.0003/10null — chance
Gemini Flash 3.8 (1st run — owner-attributed session)10/10<0.00110/10memory — subjects inside its Mar-2026 cutoff (not evidence)

Why Flash 3.8's 10/10 is not evidence

The one above-chance Round-3 run — a perfect 10/10 with formally significant p < 0.001 and textually clean rationales — is attributed to **Gemini Flash 3.8**. The owner has confirmed the run was executed in a Flash 3.8 session (a model mix-up at run time; the session was intended to be Flash 3.5). Flash 3.8's knowledge cutoff is **March 2026**, which postdates the October-2025 Nobel announcements, and two training-data channels open from that fact. The first is chart content: once the laureates became globally famous, astrology sites began publishing their charts — content that sits inside Flash 3.8's window but outside every Jan-2025 matcher's. The second is the astronomy-math channel: Flash 3.8 also knows the publicised birth dates and could, in principle, derive the four date-only signs from them. Both are training-data channels, not chart reading, and neither existed for the Jan-2025 matchers — which is exactly why the four of them scored 0/10 while this one scored 10/10. The run is the same family of memory artefact as its identical 10/10 in Round 1 on canonical celebrity charts.

Two checks confirm this is the right reading. First, the temperature-0 re-run on genuine Flash 3.5 (Jan-2025 cutoff) scored 0/10 — with the memory channel actually closed, the model reads at chance like every other Jan-2025 matcher. Second, Flash 3.8's Round-3 file is not a paste error (pairwise similarity to every other run ≤ 0.16; its rationales engage only Round-3 payload factors), and the identical rank-1 letter string it shares with its own Round-1 file is an artefact of both truth tables using the same shuffle seed (20260904) — the shuffle key is never shown to the model, and a perfect run is fully determined by the shuffle, so any perfect run prints the same letters. Re-using a seed across rounds is untidy but does not bias the model. The file is documented and archived, and set aside from the evidence because a run whose model provably holds the subjects in training is not a finding.

Round 4 — tropical-zodiac robustness probe

One legitimate challenge to any null result built on sidereal Vedic signs is: what if the zodiac itself is the problem? If the sign assignments were computed in the wrong frame, a real signal could hide behind a ~22° systematic shift. Round 4 tests this directly by re-running the Round-3 matching with everything else held fixed and only the zodiac changed: the same ten subjects, the same blind biographies, the same truth table (same shuffle seed), and the same prompt text — the word sidereal swapped for tropical — with the four factors recomputed as tropical longitudes (ayanamsa = 0, ≈22° shift, so most factors move one sign).

Why a constant shift preserves the matching problem. Adding the same ayanamsa to every planet moves every sign boundary by the same amount, so the relative structure of the ten charts is unchanged: a purely logical matcher produces the same matching in both frames. What the probe actually measures is whether the models' sign-lore differs under tropical labels — Western sign stereotypes vastly outnumber sidereal material in training data, so a tropical improvement would implicate tropical sign-lore rather than sidereal significations. Either way the result is informative: no improvement means the Round-3 null is zodiac-robust; an improvement would be a finding worth its own investigation.

Round-4 run Zodiac Exact (A) p(A) Top-3 (B) Verdict
MiniMax M3 (free)sidereal1/100.6322/10null — chance
MiniMax M3 (free)tropical0/101.0003/10null — chance
DeepSeek (browser run)sidereal0/101.0001/10null — chance
DeepSeek (browser run)tropical1/100.6325/10null — chance
Nemotron Super 120B + Ultra 550Bboth4 runs — no parseable rankings produced (format failure: models echoed the abstracts with commentary instead of emitting the required rankings)unusable — documented

Result: no tropical advantage — confirmed on two independent model families. MiniMax M3 read at pure chance in both zodiacs (1/10 sidereal, p = 0.63; 0/10 tropical, p = 1.0), and a second, much stronger instruction-follower — DeepSeek, run by hand in its browser chat with web search off — replicated the null (0/10 sidereal, p = 1.0; 1/10 tropical, p = 0.63). Pooled across families, the exact-match count is 1/20 in each zodiacal frame. DeepSeek's tropical arm does show a non-significant secondary drift (5/10 in the top 3, p = 0.15; mean true rank 4.9 vs the 5.5 null) that its sidereal arm does not share — a direction one might expect if models lean on heavily-trained Western sign stereotypes — but it never approaches the pre-registered primary bar, so it is reported as texture, not signal. The wrong-zodiac hypothesis gains no support, and the Round-3 null is not an artefact of the sidereal frame. The other free models tested could not contribute: Nemotron Super/Ultra failed the output format in all four runs (commentary and abstract-echoes, zero machine-parseable rankings — itself a data point on how little usable signal these models extract from the material), and GLM-5.2 free was rate-limited (HTTP 429) on both attempts. The Gemini matchers behind the Round-3 nulls have no free OpenRouter endpoint, so the tropical arm could not be re-run on the model family that produced the 0/10 scores.

Statistics: what the study can and cannot say

  • The null models are solidly null. Ten exact-match trials with chance expectation ≈1: 0–2/10 with p(A) ≥ 0.26 for every independent family, corroborated by shuffled-chart controls (p ≈ 0.26) and, in Round 3, by a second round on provably unmemorised charts.
  • The Gemini Round-1 scores prove memory, not reading. The probes show the memorised canonical core clears the significance bar by itself; the memory-only simulation cannot reproduce the residual; the abstract-side recognition fails. Mixed — neither cheating nor proven signal.
  • The single Round-3 10/10 is Flash 3.8's, and proves nothing. Its March-2026 cutoff includes the subjects, so the score is the same training-memory artefact as its Round-1 10/10. With it set aside, every Jan-2025-cutoff Round-3 matcher scores at chance.
  • What the study does not say. It does not say astrology's personality claims are false. It says the sign-level biography-matching signal is not strong enough for any of six independent modern language models to catch above chance — and that the famous-person matching paradigm cannot adjudicate the question at all while charts are memorised. Both are findings about the method as much as about the subject.

Honest interpretation

  • Six independent model families, zero signal. ChatGPT, Claude, DeepSeek, Grok and MiniMax all read the task exactly as instructed, passed the blinding audit, and matched at chance. Whatever these models know about astrology's sign-level personality grammar, it does not survive contact with real anonymous charts.
  • Memory is measurable and it was measured. The study's real contribution is a reusable protocol: a written protocol committed before any run, blinding audit, memory probes in both directions, a memory-only simulation bound, a shuffled-chart control, and an empirical cleanliness gate on the actual subjects before the real run. Round 1's anomaly would have been misread as evidence in any design without it.
  • The residual is intriguing but uncertifiable. Flash matched 5/5 charts it explicitly holds no confident memory of, with coherent factor-specific rationales — but on famous subjects the memory channel is structural. The rule is applied uniformly: any above-chance Gemini score on identifiable subjects is treated as memory-confounded until proven otherwise. Converting this 5/5 into evidence would require converting the Round-1 10/10s into evidence too — which the probes rule out — so it is not claimed.
  • H1 is not supported. The protocol's decision rule, fixed in writing before the runs, applied honestly across three rounds: no reproducible above-chance matching signal. Published in full, negative result included — that is the point of the programme.

Limitations

  • LLM readers, not human astrologers. A model's failure is not a human expert's failure — though it is the more objective instrument, and the one a shop customer would actually meet.
  • Sign-level factors only. Nakshatras, degrees, vargas, dashas and aspects were deliberately withheld; a finer factor set is a different experiment.
  • Date-only Round 3 has limited combinatorial spread. Several charts share Moon-in-Gemini or Venus-in-Capricorn, and four subjects carry a genuinely hour-dependent factor — discrimination is weaker than Round 1's five-factor set.
  • Sample sizes are small. Ten subjects per round is enough to show chance-level performance and to expose the memory channel; it cannot bound a small real effect tightly.
  • Same-day, exploratory — not formally pre-registered. The protocol was fixed in writing and committed to the repository before any run, but design, execution, analysis and publication all occurred on 2026-09-04. Nothing here should be read as a pre-registered trial; the next round should be registered externally (e.g. on OSF) before any data is collected.
  • The blind bios were drafted by Gemini Flash, and Gemini models were then tested on them. A real weakness, mitigated by the outcome: the non-Gemini families scored at chance on the same corpus, and Gemini's own verifiable runs scored 0/10 — a same-family fingerprint effect would be expected to inflate Gemini scores, and none appeared.
  • Temperature was not uniform across runs. Primary OpenRouter runs used temperature 0.2; the single Flash 3.5 replication used 0 — a deliberate reproducibility check after the 10/10 failed to reproduce at standard settings. Standardising to 0 for all runs would be cleaner; the null results are not sensitive to this choice.
  • Round 3 is not a replication of Round 1. The factor sets differ (five factors including lagna in Round 1; four date-only factors in Round 3), so the two rounds adjudicate nested, not identical, hypotheses. The only shared claim is the negative one: neither factor set produced a reproducible above-chance match.
  • Round 4's tropical probe covers two model families, not three. MiniMax M3 and DeepSeek both completed the task in both zodiac arms and both read at chance; the Nemotron runs failed the output format and the Gemini matchers have no free OpenRouter endpoint, so the tropical arm could not be re-run on the model family that produced the Round-3 nulls. The probe rules out a tropical advantage for the two families that could be tested; broader zodiac-robustness would need a paid or self-hosted re-run.
  • Round 3 closes the chart-content channel, not the identity channel. Matchers named the laureates 10/10 and 9/10; only chart recall was absent. The date-only design plus the null results (0/10 despite naming) contain the residual risk; it is contained, not eliminated.
  • Knowledge cutoffs were verified, not guaranteed. Cutoff dates come from vendor documentation; the cleanliness gate is the empirical backstop and passed for chart recall (identity recall is a separate, open channel — see above).

Study timeline

Date (2026) Step Outcome
09-04Protocol v1.0.0 committed to the repository; birth data verified (5 errors corrected); charts computed; bios written blindDesign fixed in writing before any run — same-day pilot
09-04Round 1 — 9 model runs (7 valid, 2 invalid-and-documented)6 families null (0–2/10); Gemini pair above bar
09-04Round 2 — §9.2 memory probes on Gemini Flash + ProVerdict: mixed, not clean — Gemini scores contaminated on the famous core
09-04Round 3 — 2025-Nobel cohort built; cleanliness gate runGate: 10/10 & 9/10 named, 0/10 chart recall (identity open, chart channel closed)
09-04Round 3 matches + replication runs4 Jan-2025-cutoff runs at chance; Flash 3.8 10/10 set aside (subjects inside its Mar-2026 cutoff)
09-05Round 4 — tropical-zodiac A/B (same cohort, same truth table, same prompt minus the zodiac word)MiniMax 1/10 vs 0/10, DeepSeek 0/10 vs 1/10 — both families at chance in both frames; Nemotron 4 runs failed the output format; no tropical advantage (pooled 1/20 each)
09-04Study closed; this article publishedH1 not supported

Further research

The protocol is the deliverable and it is reusable — and the next iteration should be formally pre-registered externally (e.g. on OSF) before any data is collected. Natural next steps: a human astrologer panel on the same Round-1 payload (to separate model capability from any signal); a Round-4 on provably astro-unpublished subjects verified empirically per matcher (the §11.4 bio-probe screen) rather than by cutoff arithmetic; finer date-only factor sets (nakshatra of the Moon, fixed by date to the day, carries more information than the sign alone) on the post-cutoff cohort; and — because the biography corpus itself can never be perfectly neutral — a control where biographies are written by multiple independent models and matched by a third. Every protocol, prompt, raw response and scoring notebook is published with this article, including the negative results, per the ATRAC Institute research programme's standing rule.

This study is scientific research on astrology's measurable claims, published by the ATRAC Institute's Sidereal Telemetry Lab. It is not a personality assessment, not a validation of any astrological product, and nothing in it should be read as a recommendation to purchase or rely on any astrological service.