Design specification: the mathematical and linguistic journey to give unique, human-sounding proper names to 55 million unnamed astronomical objects.
The International Astronomical Union has officially named fewer than 450 stars. Every other star, galaxy, pulsar, and black hole in every catalog humanity has ever compiled carries only a catalog designation — HD 189733, NGC 4889, PSR J0437−4715.
The Preservation Network holds:
| Object type | Count | Already named | Need CORE names |
|---|---|---|---|
| Stars (Milky Way) | 16,120,000 | ~450 | ~16,119,550 |
| Galaxies | 22,431,182 | ~1,000 | ~22,430,000 |
| Pulsars | ~3,300+ | ~0 | ~3,300 |
| Black holes | ~millions | ~60 | millions |
| Total | 55,000,000+ | 55,000,000+ |
The engineering challenge: generate 55 million or more names that are (a) globally unique, (b) pronounceable in any human language, (c) grand and memorable, and (d) deterministic — the same object always receives the same name regardless of when or where the engine runs.
Historical star names draw on Arabic (Aldebaran — "the follower"), Greek (Arcturus — "bear guardian"), Latin (Vega — "swooping eagle"), and Sanskrit. CORE names must feel like they belong in this tradition without borrowing directly from it.
Root material is drawn from phonological patterns across Arabic, Sanskrit, Swahili, Nahuatl, Polynesian, Norse, Latin, and Greek — giving names a cross-cultural acoustic resonance rather than a single national sound.
A name cannot begin with the same phoneme sequence as an existing named star. Regulina from Regulus, Altarina from Altair, Polarisa from Polaris — all forbidden. The names must stand alone.
The target register is the one real star names occupy: three to four syllables, open vowels, liquid consonants (L, R, M, N), a sense of weight without heaviness. Not Japanese. Not Latin scholarly text. Not computer-generated syllable strings.
Each iteration below solved the previous iteration's primary failure. The full sequence is documented here so future engineers understand why the final design is shaped the way it is.
The first engine used sixteen consonant groups (b, d, g, h, j, k, l, m, n, r, s, t, v, w, y, z) each paired with five vowels giving 80 two-character syllables. Four syllables concatenated produced an 8-character name with no parsing ambiguity.
Failed: sounded Japanese Failed: space too smallFour of the sixteen consonant groups — h, j, w, y — are the core of Japanese CV phonology. Strings like Yatanusi, Wosuramu, Janohasi were indistinguishable from transliterated Japanese. Additionally, 40.96 million names < 55 million objects needed.
Removed h, j, w, y (20 syllables). Added c, f, p series plus vowel-initial syllables drawn from real star name openings: al, ar, el, or.
Failed: echoed real star names Failed: minimal headroomThe vowel-initial group was designed to evoke existing star names — which violated the core naming rule. Altarina, Arcaluna, Sirinala all felt like derivatives. 65.6 million names covers 55 million objects but leaves almost no room for growth.
Moving to 3-character roots (CVC patterns like vel, sor, mon, kar) immediately improved name quality. Three positions at 9 characters matched the Arcturus/Fomalhaut length. With 382 roots the space clears 55 million — just enough.
Names: good Covers 55M: barely Failed: no headroom, mechanical rhythmThree same-weight positions produced names with a march-like rhythm: vel·sor·mon. No light connective phoneme between the heavy roots. With N=382 there's almost no headroom for catalog growth, and adding more roots requires curating hundreds more at consistent quality. A fourth element was needed — but not a fourth full root.
Adding a fourth root position reduced the pool requirement dramatically. For 826 million names: N⁴ ≥ 826,000,000 → N ≥ 170 roots — very achievable. Combining 2-char and 3-char roots in a prefix-free pool ensured unique parsing.
Covers 826M: yes, at N≥170 Names variable length: 8–12 chars Failed: mechanical — four heavy syllablesFour equal-weight root positions produced names with rigid, drumbeat rhythm. vel·sor·mi·bel — four hard syllables of similar weight, no natural breath in the middle. Real astronomical names have a light phoneme at their centre. This led to the binding phoneme insight of Iteration 5.
The breakthrough was recognising that the most beloved astronomical names already follow a four-part pattern with a short, open phoneme in the third position acting as a binding vowel. Position 3 is not a root — it is one of 11 vowels and diphthongs that give names their characteristic flow. This is not a design invention: it is an observation about how human languages naturally evolved these names over two millennia.
826 million unique names 15× headroom over current catalog Sounds astronomical — not generatedPositions 1, 2, and 4 draw from the same prefix-free pool of 422 curated word roots. Position 3 draws from a fixed set of 11 binding phonemes. Total name length is 7–11 characters, averaging approximately 9–10.
Real star names, examined phonetically, reveal a recurring structural pattern: a heavier onset, a secondary syllable, a light open phoneme, and a closing syllable. The light phoneme in the middle — almost always a single vowel — is what makes the name flow rather than march.
Position 3 allows five pure vowels and six classical diphthongs:
| Type | Phonemes | Examples in real names |
|---|---|---|
| Pure vowels | a e i o u | Arcturus, Belatrix, Cassiopeia |
| Diphthongs | ae ai au ia io oe | Cassiopeia, Pleiades, Hyperion |
Total unique names = P₁ × P₂ × P₃ × P₄, where P₁ = P₂ = P₄ = root pool size and P₃ = 11.
| Root pool (N) | Total space | Headroom over 55M | Verdict |
|---|---|---|---|
| 300 | 297 million | 5× | ✗ star slice alone needs 700M |
| 400 | 704 million | 12.8× | ✓ minimum viable |
| 422 (deployed) | 826 million | 15× | ✓ current pool |
| 500 | 1.375 billion | 25× | ✓ recommended for future |
Positions 1, 2, and 4 draw from the same pool. For the bijection to hold — no two distinct (P1, P2, P3, P4) tuples producing the same name string — no pool entry may be a prefix of another. If ve (2-char) is in the pool, then vel, ven, ver (3-char entries starting with ve) are excluded. Validated at engine startup.
vel + a + ri and ve + la + ri would both produce "Velari" — two different ID numbers mapping to the same name string. The prefix-free rule makes concatenation unambiguous; since names are only ever generated (never parsed back), this is sufficient.
A fixed seed (0xA57E4321) generates four permutation arrays over the pool. Each object's global rank maps through these permutations to exactly one (P1, P2, P3, P4) tuple, producing exactly one name. A scatter multiplier is applied first so that consecutive ranks produce phonetically diverse names — not a run of names with the same prefix.
# Core bijection — simplified
def id_to_name(global_rank):
scattered = (MIX_A * global_rank) % TOTAL # spread consecutive ranks
idx = scattered
p4 = idx % N; idx //= N
p3 = idx % 11; idx //= 11 # binding phoneme position
p2 = idx % N; idx //= N
p1 = idx % N
name = POOL[P1[p1]] + POOL[P2[p2]] + BINDING[P3[p3]] + POOL[P4[p4]]
return name[0].upper() + name[1:]
Roots are curated — not generated — from phonological material across 25 letter categories. Each category contributes approximately 12–20 roots: 2-char roots (usually 3 per consonant, vowels a/i/o) and 3-char roots for the remaining vowel slots.
ba, bi, bo block all 3-char roots beginning with those pairs. The remaining vowel slots (be·, bu·, br·, bl·) are available for 3-char roots: bel, ber, bra, bri, etc. Each consonant contributes 3 two-char roots and up to 17 three-char roots.
| Category | 2-char roots | Sample 3-char roots |
|---|---|---|
| B | ba, bi, bo | bel, ber, bra, bre, bri, bul, bun, bla, ble |
| C | ca, ci, co | cel, cer, cra, cre, cul, cur, cla, cle |
| D | da, di, do | del, der, dra, dre, dul, dun, dva, dve |
| E (vowel-initial) | — | ela, eli, ema, era, eri, eso, eve, eur |
| H (no 2-char) | — | hel, her, hor, hal, han, hul, hun, hir |
| K | ka, ki, ko | kel, ker, kha, khe, kra, kre, kul, kur |
| S | sa, si, so | sel, ser, sha, she, sra, sre, sul, sur |
| V | va, vi, vo | vel, ver, vra, vre, vul, vur, vla, vle |
| X (no 2-char) | — | xal, xel, xen, xil, xol, xan, xer, xur |
| Y (no 2-char) | — | yel, yer, yul, yur, yal, yar, yoa, yra |
Stars, galaxies, pulsars, and black holes live in separate database tables with separate ID sequences. To guarantee that no two objects of any type ever share a name, CORE assigns each type a global offset: a reserved, non-overlapping slice of the 826-million-name space.
Stars get 700 million slots — 84.7% of the entire name space. This is deliberate. CORE names Milky Way stars only. The stars of other galaxies number in the hundreds of trillions; even if every one were catalogued, naming them individually would be an undertaking of an entirely different order. The Milky Way itself contains an estimated 200–400 billion stars. 700 million slots covers the full realistic scope of future Gaia-era surveys with room to spare.
| Object type | Offset start | Slice | % of space | Rationale |
|---|---|---|---|---|
| Stars | 0 | 700,000,000 | 84.7% | Milky Way — present + all future surveys |
| Galaxies | 700,000,000 | 100,000,000 | 12.1% | Current 22.4M + deep-field growth |
| Pulsars | 800,000,000 | 10,000,000 | 1.2% | ~3,300 known + future detections |
| Black holes | 810,000,000 | 10,000,000 | 1.2% | Confirmed + stellar candidates |
| Other objects | 820,000,000 | 6,000,000 | 0.7% | Neutron stars, white dwarfs, etc. |
| Reserved | 826,000,000 | 665,928 | — | Rounding buffer to TOTAL |
Each object's global rank = its offset + its sorted position within its type. This rank feeds directly into the bijection. The resulting name is globally unique: no star and galaxy, however distant from each other in the sky or the database, will ever share a name.
The fixed seed (_SEED = 0xA57E4321) must never be changed. The scatter multiplier (_MIX_A = 316,227,767) is computed at startup from the pool size and is also stable as long as the pool is unchanged. Adding new roots to the pool would shift every name and must be treated as a breaking change requiring a full re-run across all tables.
The following are real output from the deployed CORE engine — not hand-crafted examples. Each name is the deterministic result of its object's global rank passing through the bijection.