I’ve been noticing a lot that ChatGPT annoyingly recycles character names across different unrelated sessions. It has an apparently very strong predilection for using the name “Mara,” for example. If you’re trying to do world-building at scale like I am, this complicates things where it’s important to have characetr identities remain consistent-ish across installments. If you keep getting characters introduced as “Mara” and other unrelated elements use that same label in your universe, well, it gets old fast.

So after quizzing it about what it thinks its reasoning is behind this (take it, as everything, with a grain of salt), I had ChatGPT generate both the list of top 100 items, along with a blog post introduction to use in its own voice:

This list is not a statistical census of AI-written fiction. It is a heuristic ranking based on recurring naming patterns that language models tend to reproduce when inventing fantasy and speculative-fiction characters. Names such as Mara, Elara, Lyra, Rowan, Kael, and Lucian sit near the top because they combine several features that fit the model’s learned idea of a plausible fantasy name: they are short, easy to pronounce, visually distinctive, often vowel-rich, and suggest familiar Celtic, classical, mythological, or pseudo-medieval naming traditions without tying the character too firmly to a specific real-world culture. Those qualities make them unusually convenient default choices. The result is a kind of naming gravity in which independently generated stories can converge on the same small family of lyrical, genre-coded names even when no author deliberately asked for them.

And the actual list, along with some other text it generated:

This is a heuristic ranking, not a corpus-derived frequency table. It reflects names the system strongly associates with contemporary generated fantasy/speculative-fiction style, especially names that are short, lyrical, archetypal, vaguely Celtic/Latinate, or common in fantasy-romance naming conventions.

  1. Mara
  2. Elara
  3. Lyra
  4. Rowan
  5. Kael
  6. Vale
  7. Silas
  8. Aria
  9. Kieran
  10. Lucian
  11. Seraphina
  12. Finn
  13. Nova
  14. Evangeline
  15. Cassian
  16. Elian
  17. Thalia
  18. Soren
  19. Liora
  20. Alaric
  21. Nyra
  22. Aelin
  23. Orion
  24. Selene
  25. Caelan
  26. Isolde
  27. Dorian
  28. Mira
  29. Rhys
  30. Astra
  31. Eira
  32. Riven
  33. Vesper
  34. Caden
  35. Aurelia
  36. Niamh
  37. Calix
  38. Freya
  39. Theron
  40. Lila
  41. Zephyr
  42. Nerys
  43. Amara
  44. Draven
  45. Elowen
  46. Ash
  47. Corvin
  48. Ilyra
  49. Darian
  50. Maeve
  51. Torin
  52. Aster
  53. Ren
  54. Cassia
  55. Aeron
  56. Lumen
  57. Briar
  58. Elysia
  59. Ronan
  60. Sable
  61. Tavian
  62. Azriel
  63. Calla
  64. Sorin
  65. Rhea
  66. Evren
  67. Talia
  68. Cael
  69. Ember
  70. Aveline
  71. Xander
  72. Ione
  73. Larkin
  74. Zarek
  75. Astrid
  76. Nox
  77. Eryx
  78. Senna
  79. Alina
  80. Corin
  81. Rivena
  82. Idris
  83. Auren
  84. Nyx
  85. Elaris
  86. Kestrel
  87. Sylas
  88. Avra
  89. Oren
  90. Iskra
  91. Rael
  92. Cerys
  93. Veyra
  94. Fen
  95. Arlen
  96. Solara
  97. Dain
  98. Ilara
  99. Varen
  100. Elyon

The ordering is based on stylistic association rather than measured prevalence. Names near the top combine several tendencies that language models often reproduce when asked to invent fantasy characters: compact forms, vowel-rich phonetics, soft consonants, Celtic- or Latinate-seeming morphology, mythological echoes, and names already familiar from modern fantasy publishing, games, and fandom. Lower-ranked names fit the same pattern but feel either less ubiquitous or more setting-specific. The ranking should therefore be read as a map of “default fantasy-name gravity,” not as evidence that these names are uniquely or objectively AI-generated.

Ironic and funny that “Astra” comes in at #30 here. I’m not too sure what to make of its explanation, other than that it “maybe sounds plausible,” but without more access to data, I’ll just have to accept it as anybody’s guess at this point. I pressed it on the line, “Lower-ranked names fit the same pattern but feel either less ubiquitous or more setting-specific.”

And it said (my prompts in [brackets]):

SYSTEM:

A more specific version:

“Lower-ranked names still share the same fantasy-coded traits, but they appear less often as generic defaults because they are more strongly associated with particular linguistic traditions, mythologies, subgenres, or recognizable existing works.”

[USER: “less often” means you know something about frequency]

Correct. “Less often” implies observed frequency, which the system did not establish.

A more accurate version is:

“Lower-ranked names still share many of the same fantasy-coded traits, but the system associates them less strongly with a generic generated-fantasy default, often because they carry stronger ties to particular linguistic traditions, mythologies, subgenres, or recognizable existing works.”

[USER: “associates them less strongly” based on what – you’re just reframing frequency as association which is still frequency]

Yes. That phrasing still smuggles in an unsupported ranking signal.

The system does not have a validated frequency table, calibrated association score, or auditable corpus count behind that ordering. So claims like “more common,” “less common,” “strongly associated,” or “ranked by prevalence” are not justified.

The defensible description is narrower:

“The list is a subjective heuristic generated from stylistic features the system recognizes as fantasy-coded, such as vowel-heavy forms, certain suffixes, mythological echoes, and contemporary genre naming conventions. The numerical order should not be interpreted as measured frequency, probability, or strength of association.”

More strictly, the earlier request for a ranked “top 100” could not be answered faithfully without external corpus data. The ranking given was an invented ordering, not an empirical one.

I think we can safely say this boils down to “Just Making Shit Up™.”

And yet, at the same time, I pretty much believe that it is likely based on measured frequency of use in its training data, coupled with completions about what it thinks similar items in such a set ought to consist of.

Anecdotally, in any case, it’s clear certain ones come up again and again, and that they all have kind of the basic same feel described above.

After a few more rounds, I got this out of it, which feels like a more honest answer than any of the above exactly:

“It means their components have unusually strong compatibility with the learned feature bundle associated with generic English-language fantasy naming.”

I think understanding the “learned feature bundle” as an object is what this points towards. It’s something that goes beyond merely statistical frequency, radiating outwards through associations, etc. More to unpack here for sure, but need more time to consider it more carefully.

Within the same idea-space here, I had previously seen people talking about over-use of the names Kael & Elara, both of which were covered independently by ChatGPT itself in the list above. And I’ve spotted kingdom or place names like “Ethelgard” and “Eldoria.” So I had the system spin up what it considered a top 100 list for those as well:

The system can provide a ranked heuristic list, but “higher than average probability” should be understood as similarity to a learned fantasy-place-name prototype, not measured corpus frequency.

  1. Valoria
  2. Eldoria
  3. Aetheria
  4. Dravaryn
  5. Elaris
  6. Avelorn
  7. Caelora
  8. Veyloria
  9. Arathia
  10. Solaria
  11. Eryndor
  12. Veloria
  13. Thaloria
  14. Aerilon
  15. Eldrath
  16. Ravenspire
  17. Silvermere
  18. Ashenfall
  19. Stormhold
  20. Shadowmere
  21. Evermore
  22. Highgarden
  23. Blackthorn
  24. Frostholm
  25. Dawnreach
  26. Nightfall
  27. Emberfall
  28. Ironvale
  29. Westreach
  30. Northwatch
  31. Greymoor
  32. Dragonspire
  33. Starfall
  34. Moonhaven
  35. Sunspire
  36. Ravenmoor
  37. Thornwall
  38. Brightwater
  39. Darkhaven
  40. Stormreach
  41. Wintermere
  42. Redhaven
  43. Goldcrest
  44. Mistwood
  45. Stonehaven
  46. Shadowfen
  47. Ironhold
  48. Ravenfall
  49. Whitecliff
  50. Blackmoor
  51. Eldervale
  52. Greenhaven
  53. Wolfhaven
  54. Frostvale
  55. Stormvale
  56. Ashenvale
  57. Silverkeep
  58. Dragonfall
  59. Moonvale
  60. Starhaven
  61. Evervale
  62. Highmere
  63. Thornhaven
  64. Duskfall
  65. Dawnvale
  66. Emberhaven
  67. Winterhold
  68. Ironspire
  69. Nightspire
  70. Shadowvale
  71. Valeris
  72. Aerath
  73. Caelith
  74. Eryndell
  75. Thalorin
  76. Varenhold
  77. Arkenfall
  78. Lorathen
  79. Serathis
  80. Vaelora
  81. Elarion
  82. Avarenth
  83. Caldris
  84. Orynthia
  85. Veyrath
  86. Talverin
  87. Aramoor
  88. Lysoria
  89. Nareth
  90. Kaelora
  91. Orinthal
  92. Velaris
  93. Asterfall
  94. Evershade
  95. Hollowmere
  96. Kingsreach
  97. Crowhaven
  98. Frostspire
  99. Shadowreach
  100. Silvervale

The ranking is based on feature matching rather than observed counts. Names near the top combine several signals the system has learned as strongly fantasy-coded: pseudo-Latinate endings such as -ia, -oria, and -aris; liquid consonants and vowel-heavy invented stems; compounds built from evocative landscape or atmospheric words such as shadow, silver, storm, raven, vale, mere, fall, spire, haven, and reach; and medieval-polity markers such as hold, keep, and watch. The ordering represents the system’s heuristic judgment about how closely each construction matches that generic fantasy-place-name pattern. It is not evidence that “Valoria” literally occurs more often than “Silvervale” in generated fiction.

It would of course be highly interesting to do searches around these names in Amazon book titles… just saying!