Stars Preview
The full stellar catalog browse by spectral class, magnitude, or name.
The full stellar catalog browse by spectral class, magnitude, or name.
With over 55 Million+ astronomical objects across the entire catalog (stars, galaxies, quasars, variable stars, exoplanets, pulsars, and much more) writing a genuinely distinct description for each one was not a task any single human could handle without going insane, even if you enjoy doing it, which we found that no one did. The task of producing millions of paragraphs of content to satisfy our internal pledge of having the cleanest, most complete content for the STEN and GLEN engines meant that some work in automation and parallelism would be needed. We then thought about using AI to accomplish this task.
However, we soon discovered that even with AI and LLM, including a trial of setting up our own model (Ollama) to be trained to run this massive job, we soon realized that many of the millions of stars would have said or shared the same sentence. Artificial Intelligence currently lacks context, and even with the greatest of prompts, this would have been a futile effort, and mostly all content would have shared the same paragraph word for word as a mathematical and statistical certainty.
So to solve the problem of variableness and uniqueness across over 50M+ entries, we built a dedicated description engine for each of the object types in the entire catalog. Each engine assigns an archetype based on the object’s own measured physics, not randomly. A quasar at redshift z > 3.5 is classified as an Ancient Fire; one from cosmic noon is an Epoch Peak; a nearby ultrabright source becomes a Nearby Giant. A variable star with period over 100 days is a Long Breather; a Cepheid becomes a Rhythm Keeper; an eclipsing binary star is an Eclipse Dancer.
This appears to be simple and rudimentary by nature, but it is what allowed the Description Engine to create variableness for the massive job of writing stories and descriptions for millions of stars with no name, and the millions upon millions of objects occupying our night sky. Within each archetype, a SHA-256 hash seeded by the object’s unique catalog ID then selects phrases from large independent pools: an opening hook, several content modules each expressing physical data in varied phrasing (period, amplitude, distance, constellation, spectral class, flux, orbital data), and a closing line.
The Description Engine works beautifully. To wit, we have been able to run it as a local script on a basic PC, and achieve an output of over 3,000 descriptions per second for mostly all stars, with the 6M+ Quasar category alone completing in a mere 33 minutes, all with unique and different descriptions for each one. This was amazing to see. To this end, the engine works so successfully that the same object always produces the same description, and a different object will always produce a different one, even within the same archetype.
Taking just ten archetypes multiplied by fifteen phrase variants across eight independent selection modules yields well over 25 billion base combinations per object type, which is roughly 10,000 times larger than the biggest single peer group (2.3 million long-period variable stars). Adding module ordering variation and data-interpolation phrasing pushes the theoretical space into the hundreds of trillions. No two objects of the same physical class need share a single sentence, and the task of writing and generating tens of millions of unique descriptions for the various object types in the catalog was achieved without the impossible task of examining each star’s specifications and writing by hand, its description and story.