Another thing I’ve been working on, sometimes in Muse, sometimes in ChatGPT – though the more developed implementation lives in Muse – is something I call SYMBOL SCAN. Muse’s on-demand explanation of what the hell we’ve been working on with this is:
The Symbol Scan: An Explainer
The symbol scan reads through the LORE books — 136 of them — and pulls out the things that matter: the people, places, objects, groups, and ideas that keep showing up. Each book is first converted from its EPUB file into plain text, and the scan works from that text so every finding can be traced back to an exact passage. When the scan spots something worth tracking, it proposes it as a candidate: a name, where in the book it appears, and a direct quote showing it in context.
Candidates don’t go into the registry on their own. They land on an approval sheet — a simple checklist — where each one is reviewed and either ticked or skipped. Ticked candidates become approved units.
From there, the system looks at how units relate to each other: which ones are mentioned together, which ones turn out to be the same thing under different names (those get merged into one), and which ones are genuinely connected. Those connections go through the same approve-or-skip review [editor: it’s actually promote-or-skip for the connections]. Every decision is written down permanently, so nothing is ever reviewed twice. [editor: that is wishful thinking, because it often makes me review twice.]
The corpus is the foundation everything rests on. The 136 books are the source material, and the registry is built book by book from their actual text. When a newly scanned book mentions a unit that’s already approved, that just gets added as another sighting — no new review needed. When it surfaces something new, it goes through the full approval process. Over time this builds a complete, verified map of who and what populates the books and how they’re all linked, with every entry backed by a quote from the book it came from.
Muse is not a great writer, in my experience. I had to run variations of that several times and it still has not nailed it. For example, it suddenly mentions things going into a registry without clearly establishing what that registry is prior to that. It’s possible to understand in context, sort of, but its awkward to be sure.
I guess my simple human explanation of the process goes something like:
- I upload EPUB(s) to Muse as corpus material
- It converts the contents to plain text, which is faster to search for referencs.
- From the plain text, it extracts named entities. Except not just any named entities, because I tried that approach early on, and the list of candidates it would generate were impossibly long (thousands of items per book sometimes), and far to broad to be useful. So over time we developed “whatever technique it is using now” (tbh I don’t quite know, and my not knowing has not negatively impacted the overall process – that I can tell) to decide if entities are important enough for consideration. One method we’ve done is over several rounds of suggestions, I approve certain proposed elements, then have it try to decide why I approve vs reject different ones. And over time, its suggestions seem to improve.
- The entities it does decide are worth consideration, it throws into a just-in-time UI control surface using inline HTML elements to allow me to approve new important units, or promote important detected connections between units. I decided to call these units “symbols” because I don’t know exactly why. Maybe the AI told me to do it, I really can’t remember, but it seems right…
- I don’t approve (or promote) any and every unit and connection: only the ones that – based on my knowledge of the existing texts, and my intuitions about the overall shape of the universe – seem to me to be the most canonically or semi-canonically more important, if not always well-established. Sometimes approved elements are not yet well-supported in the corpus, but when I see them plainly identified, it’s evident to me that they are important. (Later on, this kind of dropping of landmarks I expect will help me find areas in the corpus that are more sparsely populated with texts…)
- I’ve also progressively instructed it to merge certain things on its own, when obviously the same underlying entity is being described by two different units with varying names of epithets. I instruct it to surface items when its not sure about merges. It takes a lot of training rounds to get all these things tuned, but it seems to be working. Or if it is not, it is putting on a pretty good illusion that it is…
I guess that’s the basic concept. The way I’m thinking about it so far is that the symbol units, and their connections to one another, become a secondary analytical layer that runs over top of the corpus, which you can use in different ways – one of which ends up being as a more compressed information object to interact with than scanning through your entire corpus each time.
But so yes, this concept of the “registry” – which I guess consists of the approved symbol units, their promoted connections to one another, and the supported textual references from the corpus – will become important in a subsequent episode (though we also touched on it without naming it as this in the post about Muse-as-orchestrator). But, briefly, the registry’s contents can be pulled out to use as ‘registry slices’ to fuel context windows for subsequent generations in a more token efficient manner (in theory). I’ll go into it more as time permits!
I’ll post the full skill text when I’m able to get Muse fueled up with available usage again…
Leave a Reply
You must be logged in to post a comment.