Tim Boucher

Questionable content, possibly linked

Visualizations of AI Authoring app: tape.deck

That draft of the simple spec for tape.deck that 6.1 Sol generated is “pretty good.” Not perfect, but pretty good. Some things I would have excluded for initial tighter focus, but also, whatever. Some of it represents a sort intellectual archaeology of how the idea and the process have evolved too, so don’t want to cut too much of that out. But neither do I want to overwhelm potential readers with TMI, which is incredibly easy to do when your workstream involves generative AI as an essential part of your tooling and ideation.

So, I thought a next useful step would be expose another side of those explorations that I’ve done to show some of the ground that has been covered, and what the findings were, so as to bring it all back to the main branch eventually to continue composing pieces of the amalgamated toolbox that tape.deck represents.

Anyway, what follows are simplified wireframes I had ChatGPT do of different key screens within what I imagine one implementation of tape.deck could look or function like:


Story Space view

So this screen visualizes a theoretical corpus – in my case it is EPUB files from a linked narrative multiverse project. And this would be generated after subjecting those files to conversion to plain text formats, and then running my SYMBOL SCAN skill over them to pull out the most important-seeming entities (“symbols”) and associations between them for review and promotion.

This story space view then lets you visualize those various symbols and their connections as linked clusters, which can then be expanded or recombined in other ways in other views.

Registry view

For each symbol node that appears in the story space, there is a corresponding registry entry. This is a record of all known information about a given symbol, including its “attestations” – which books it appears in, what quotes reference it – key links to other symbols, usage rules or restrictions for that symbol.

[Apparently Lucasfilm has a FileMaker database they use for a similar purpose for Star Wars in-universe continuity, which is called The Holocron.]

As the composition process progresses through the workspace, one of the things that gets included in the capsules for any finished “cassette” (to be sent out to renderers or playback environments) is registry slices. That is, one of the key capsules that gets transmitted with cassettes is enough information from the included symbols’ registry entries to allow renderers to faithfully reproduce the underlying cassette’s contents.

Builder view

These next few screens feel a bit jumbled to me, and they have some different tab or section labeling that is inconsistent with each other. But the point this is trying to make is that cassettes can be composed out of nested units: capsules (and sub-capsules), blocks, frames. These can be chained together as sequences, or linked via decision-points for branching narrative outcomes. The above screenshot is intending (I think) to show the ability to arrange and edit capsules, chains, and blocks together in a Builder interface.

Anyway, these are all just initial visualizations and the specifics are not all completely nailed down yet.

Cassette view & preview

Two different views into the same finished cassette, the first showing more the complete package (capsule flow, registry, etc), and the second a more focused view into the specific linear content elements in their arranged final forder:

Focused version:

Rendered outputs view

From the finished cassette, various types of rendered outputs could be produced, either by the creator of the cassette directly themselves, or by third party renderers or playback environments which might read from the cassettes down the road. The most obvious modalities that come to mind for rendered outputs would be text, audio (narration, music, sound effects), video, VR artifacts or scenes, etc.

As I wrote about in my original post about a workflow for using AI in video, each of the generated outputs is pegged to the frame, block, or capsule which it depicts or renders through available media formats. Conceivably, the creator of a cassette (or a cassette’s own automated evidence gate checks) could go through a review queue of generated outputs and approve or reject items which do or do not match the ideal vision represented by the cassette. And then, in those cases, approved artifacts could be included with the cassette as it travels onwards to other generation or playback environments.

Or presumably, from there, the cassette creator could simply output various different complete renders: a finished ebook (or series), a dramatised audio book, a complete video, etc. And then the selected artifacts made available are the only authorized ones by the cassette’s creator, and the underlying cassette does not travel. And they could always come back in a few years and use new technologies to re-render the underlying story structure without having to re-write it.

Anyway, hopefully the visualizations help ground it more!

tape.deck: A Medium-Independent Authoring & Composition Tool

It’s difficult to describe the process behind using AI to build a tool that you don’t quite know what it is, but now that you have the ability to do it, you’re compelled to continue and to find out.

These things sometimes take time to reveal themselves though, and I think I’m starting to understand what this one is and how it functions, out of the other bits and bobs I’ve managed to piece together of the “grand idea” (more of an intuition) that is fueling the whole thing.

For now, it’s called tape.deck.

Here’s a faux diagram that I got ChatGPT to make, and I think grounding it in a physicalization really helps for comprehension:

The idea in a nutshell is this:

Generative AI makes it possible to realize the same composition in many forms. This calls for tools that let authors shape a story’s underlying structure and guide how it becomes text, images, audio, or video.


After much back and forth, I managed to coax a pretty coherent simple specification for what this app/system/workflow/pipeline/etc. ought to be able to perform – at least for now. On-going usage and exploration would guide future development. Included below is GPT-6.1 Sol text:


tape.deck — simple app specification

Version 3 · October 8, 2026

What the app does

tape.deck supports medium-independent authoring and composition. The author arranges symbols, blocks, capsules, and instructions into a cassette that different renderers can realize as text, images, narration, or video. Audio and visual references can guide those realizations.

Author once, render many.

A cassette contains the underlying composition and the information needed to interpret it. Different treatments can change its expression while preserving its authored arrangement and constraints. Specific human-authored text can be included when exact wording matters.

The composition package

  • Cassette — the whole package from which a medium-independent composition can be mixed and rendered.
  • Head capsule — the first capsule; describes the cassette and holds shared instructions, mix settings, and rendering constraints.
  • Registry slices capsule — contains background references, definitions, and allowed usages for the symbols included in the composition.
  • Content capsules — flexible containers holding blocks or nested sub-capsules. Capsules can be chained together in sequences.
  • Blocks — groups of compositional material that can be chained into sequences within a capsule.
  • Frames — sit inside blocks and identify which symbols perform which actions upon which other symbols.
  • Symbols, or symbol units — addressable elements drawn from the registry, including people, places, objects, groups, and ideas.

Setting and instructions can be established at cassette, capsule, or block level. Descendant material inherits that context, with explicit local refinements. Conflicting constraints remain visible for resolution.

Main functions

Story Space

Browse the fictional universe through its source material, symbols, and relationships. A searchable network opens the supporting passages and recorded connections. Multiple sourced versions can coexist. Later, the app can suggest promising paths or unexplored connections as starting points for new compositions.

Registry

Import books and scan them for meaningful symbols. Proposed entries, relationships, and merges go through human review. Approved entries retain their source passages, descriptions, connections, and usage rules. Audio, images, and other files can be attached as references. Review decisions persist.

Builder

Compose frames within blocks, arrange blocks within capsules, and sequence capsules into a cassette. Support nesting and occasional branches through a zoomable flow. Selecting an element opens its composition and instructions; relevant registry material appears contextually. Direct text entry supports guidance and exact inclusions.

Mixing

The head capsule holds mix settings: desired treatments, constraints, and allowable renderer choices. These can include tone, emphasis, viewpoint, permitted invention, language, medium, style, format, and genre.

Renderers apply those settings through their own capabilities and may offer further adjustments within the cassette’s constraints. A revised mix can produce another realization while preserving the composition and earlier results.

Preview

Inspect the assembled cassette before rendering. Preview is a focused, read-only document over a dimmed or blurred background, with collapsible navigation and a compact flow overview. It shows the package and effective instructions being sent to the renderer. Registry details open when selected.

Rendering and results

Run the cassette through different renderers and retain the resulting versions. Read text, play narration and video, and open images large with arrow-key navigation. Compare alternatives, approve or reject results, revise mixes, and export selected outputs.

Playback and distribution

The distributed object is the cassette. A network, channel, distributor, or other recipient can purchase or license it and produce its own realization for an audience or subscribers [or for personal use]. The underlying composition persists across those versions.

Playback is the experience-facing operation. It can present an existing render or generate a realization as it proceeds. A playback environment interprets the cassette’s authored arrangement, symbol references, allowed usages, and mix settings.

An interactive fiction experience could let audience choices influence the realization within permitted boundaries. A device could realize a cassette locally, or present a version produced elsewhere. These are possible applications of the same package and rendering model.

Manifest and review

Each rendering includes a manifest describing the basis of its content decisions and additions made by the renderer. Required human-authored wording can be pinned to persist verbatim in applicable outputs.

A separate review checks the realization against the cassette and flags omissions, undeclared additions, instruction drift, or unexpected symbol usages. Useful new usages can be approved and added to the registry. Uncertain findings remain available for human judgment.

The output can travel with a generated bill of materials, or GBOM, containing its input references, manifest, and review record.

Curation and reuse

Retain approvals, rejections, edits, and review decisions. Selected finished work can be scanned back into the registry, with new material proposed for approval. Over time, this record can support better suggestions and an experimental encoding of the author’s preferences.

Shared behavior

Work and decisions survive closing the app. Editing and browsing do not require model calls. Generation is an explicit action. Earlier compositions, mixes, and outputs remain available, and cassettes can be exported for use in other rendering or playback environments.

The main views are Story Space, Registry, Builder, Cassette, focused Preview, and Render Results. Mixing belongs within the Builder and cassette configuration. The visual style remains open.

First complete version

The first version supports importing source books, reviewing registry entries, composing and sequencing capsules and blocks, packaging and previewing cassettes, rendering text, reviewing results, and saving or exporting the work.

Images, narration, and video extend the same structure. Distributor playback, interactive experiences, suggested paths, taste encoding, and Content Credentials integration can develop from that foundation.


I’ll continue this in a separate post or this is going to get too long to parse all at once…

First Reader

I’ll expand on this another time, but wanted to jot down this idea that with the lorecore books, I am not possessive about them from an “authorial sense.” Did I “write” them? Some yes, very much, some less. But there are myriad choices, branches, possibilities, taken or untaken. Much of it is merely the result and record of what is explored, discovered, provoked, or simply noticed. I think of myself more than anything as being the First Reader of one of these new works. And then I apply my own sort of human-intelligence filtering workflows against it until it matches the criteria of quality that I am looking for at any given moment (which fluctuates with time and further experimentation).

Building the Lorecore Registry

Okay, so I’ve topped up my AI subscriptions for the month, and got Muse to shoot out a kind of coherent statement about what happens after the SYMBOL SCAN skill explained earlier, in relation to my AI Lorecore books.

Basically, what happens is that once books are converted and scanned, the process pulls out working narrative units, “symbols,” which taken together with their metadata form a “registry” which Muse explains adequately in the passage below:

The registry is a catalog of everything that matters across the books. Its basic unit is the symbol: a person, place, object, group, or idea worth tracking. Each entry holds the symbol’s name, the passages where it shows up, and the other symbols it connects to, all backed by quotes from the text. It can be read in both directions — look up what the copper bit is, or look up which symbols appear in book 12.

It gets built by a process called the symbol scan. Each book is converted from its EPUB into plain text, which is faster to search and lets every finding point back to an exact passage. The scan reads through and proposes candidates: a name, where it appears, a quote showing it in context. Nothing enters on its own. Every candidate lands on an approval sheet — a simple checklist — where each one is looked at and either approved or skipped. Approved symbols become entries.

The links between symbols go through the same review. When two symbols keep appearing together, the scan proposes a connection; when two names turn out to be the same thing, it proposes a merge. Each proposal is either promoted or skipped, and the decision is written down permanently, so the system keeps a record of what has been decided. When later books are scanned, sightings of known symbols are added to their existing entries, and only genuinely new material goes through review.

Once the registry exists, it can be sliced — a subset of entries pulled out for a particular job — and entries can be linked together in new ways. The catalog is the raw material; a slice is the working set.

We’ll get into later how registry slices can be used, and talk about the mechanics of using and re-using symbol registry entries…

Michael Smith Streaming Fraud Sentencing

I am embarrassed to read this kind of “human slop” describing the sentencing of Michael Smith, who ran an automated artificial streaming ring to the tune of $8M:

“Michael Smith exploited super intelligence technology to generate a fraud,” said U.S. Attorney Jamie McDonald. “By flooding music streaming platforms with automated bots in the place of consumers, and fake songs in the place of creativity, Smith robbed millions in royalty payments from genuine artists and their fans. This Office is committed to ensuring the integrity of all markets, and protecting the public from those who use super intelligence for fraud.”

My understanding is that, Executive Orders notwithstanding, “super-intelligence” as a technology simply does not exist yet (let alone absolutely did not between the time period his scheme was operational, 2017-2024), making the statements above verifiably factually inaccurate.

Was there automation? Sure seems like it, based on sources I’ve seen. This one appears to give a more complete run-down of what actually happened, without inserting any pro-American technology propaganda:

In other words, Smith’s fraudulent activities involved misrepresenting information to the streaming platforms, creating false accounts, and disguising the true nature of the streams, using bot accounts rather than human listeners, according to the indictment.

Smith used software to continuously stream songs he owned, according to the indictment. He also allegedly paid co-conspirators and people overseas to sign up for bot accounts. Smith used false names to sign up bot accounts and used debit cards in fake names to pay for the accounts, the indictment stated. […]

To avoid detection, Smith spread the artificial streams across tens of thousands of songs, where each streamed a smaller number of times to appear more credible. At its peak, Smith generated about 661,440 streams per day, resulting in annual royalties to himself exceeding $1.2 million, according to the indictment. […]

To fuel the operation, Smith collaborated with an AI music company and a music promoter to create hundreds of thousands of AI-generated songs. The music company and promoter are not named in the indictment.

But, simply including automation – a.k.a. “software” – does not automatically make everything it touches into “super-intelligence.” Was what he did “intelligent”? That could be argued either way, but it certainly takes a lot of effort and orchestration to do what he did, at such a scale over such a long time. And in a way, this kind of fraud at its root is really just hyper-capitalism, trying to directly hack the reward mechanism without adding anything of value to society in exchange. Following through on that line of thinking, it makes a certain amount of sense to say that the market is already the “super-intelligent” paperclip maximizer people talk about as though it’s some far off possibility. We’re already enslaved to a machine which is trying to convert every resource, moment, and interaction into dollars. I don’t condone what Smith did here, of course, but it’s a logical endpoint of the games we are all forced to play all day everyday online and offline (if offline even exists anymore as a separate thing). And it is a direction which is only going to get more and more insane and exaggerated as AI ramps up…

Symbol scan description

Another thing I’ve been working on, sometimes in Muse, sometimes in ChatGPT – though the more developed implementation lives in Muse – is something I call SYMBOL SCAN. Muse’s on-demand explanation of what the hell we’ve been working on with this is:

The Symbol Scan: An Explainer

The symbol scan reads through the LORE books — 136 of them — and pulls out the things that matter: the people, places, objects, groups, and ideas that keep showing up. Each book is first converted from its EPUB file into plain text, and the scan works from that text so every finding can be traced back to an exact passage. When the scan spots something worth tracking, it proposes it as a candidate: a name, where in the book it appears, and a direct quote showing it in context.

Candidates don’t go into the registry on their own. They land on an approval sheet — a simple checklist — where each one is reviewed and either ticked or skipped. Ticked candidates become approved units.

From there, the system looks at how units relate to each other: which ones are mentioned together, which ones turn out to be the same thing under different names (those get merged into one), and which ones are genuinely connected. Those connections go through the same approve-or-skip review [editor: it’s actually promote-or-skip for the connections]. Every decision is written down permanently, so nothing is ever reviewed twice. [editor: that is wishful thinking, because it often makes me review twice.]

The corpus is the foundation everything rests on. The 136 books are the source material, and the registry is built book by book from their actual text. When a newly scanned book mentions a unit that’s already approved, that just gets added as another sighting — no new review needed. When it surfaces something new, it goes through the full approval process. Over time this builds a complete, verified map of who and what populates the books and how they’re all linked, with every entry backed by a quote from the book it came from.

Muse is not a great writer, in my experience. I had to run variations of that several times and it still has not nailed it. For example, it suddenly mentions things going into a registry without clearly establishing what that registry is prior to that. It’s possible to understand in context, sort of, but its awkward to be sure.

I guess my simple human explanation of the process goes something like:

  1. I upload EPUB(s) to Muse as corpus material
  2. It converts the contents to plain text, which is faster to search for referencs.
  3. From the plain text, it extracts named entities. Except not just any named entities, because I tried that approach early on, and the list of candidates it would generate were impossibly long (thousands of items per book sometimes), and far to broad to be useful. So over time we developed “whatever technique it is using now” (tbh I don’t quite know, and my not knowing has not negatively impacted the overall process – that I can tell) to decide if entities are important enough for consideration. One method we’ve done is over several rounds of suggestions, I approve certain proposed elements, then have it try to decide why I approve vs reject different ones. And over time, its suggestions seem to improve.
  4. The entities it does decide are worth consideration, it throws into a just-in-time UI control surface using inline HTML elements to allow me to approve new important units, or promote important detected connections between units. I decided to call these units “symbols” because I don’t know exactly why. Maybe the AI told me to do it, I really can’t remember, but it seems right…
  5. I don’t approve (or promote) any and every unit and connection: only the ones that – based on my knowledge of the existing texts, and my intuitions about the overall shape of the universe – seem to me to be the most canonically or semi-canonically more important, if not always well-established. Sometimes approved elements are not yet well-supported in the corpus, but when I see them plainly identified, it’s evident to me that they are important. (Later on, this kind of dropping of landmarks I expect will help me find areas in the corpus that are more sparsely populated with texts…)
  6. I’ve also progressively instructed it to merge certain things on its own, when obviously the same underlying entity is being described by two different units with varying names of epithets. I instruct it to surface items when its not sure about merges. It takes a lot of training rounds to get all these things tuned, but it seems to be working. Or if it is not, it is putting on a pretty good illusion that it is…

I guess that’s the basic concept. The way I’m thinking about it so far is that the symbol units, and their connections to one another, become a secondary analytical layer that runs over top of the corpus, which you can use in different ways – one of which ends up being as a more compressed information object to interact with than scanning through your entire corpus each time.

But so yes, this concept of the “registry” – which I guess consists of the approved symbol units, their promoted connections to one another, and the supported textual references from the corpus – will become important in a subsequent episode (though we also touched on it without naming it as this in the post about Muse-as-orchestrator). But, briefly, the registry’s contents can be pulled out to use as ‘registry slices’ to fuel context windows for subsequent generations in a more token efficient manner (in theory). I’ll go into it more as time permits!

I’ll post the full skill text when I’m able to get Muse fueled up with available usage again…

Disabling backend workers in Muse

Regarding my last post about Muse routinely wiping and rebuilding itself for some reason, it finally replied to my inquiry as to why it was blowing up my usage with:

The backend worker runs every 2 minutes and the panel poller every 3 minutes; both are generating usage with no pending work. The system can reduce or stop them.

Options:

  • Disable both background checks (generation requests and panel decisions would only be handled when requested directly).
  • Slow them to every 30 minutes.
  • Keep them running as-is.

And then it offered some corresponding buttons for the options. I picked disable, but it’s difficult to trust these systems to be doing what they say they’re doing, or if you click a button like this if it will have the effect it says, or what the unintended consequence of making this selection will even be?

Totally unclear for now, but guess we’ll find out.

After it again reminding me my usage is out, it confirmed:

Both background checks are disabled. No further scheduled runs will occur unless re-enabled by explicit instruction.

Note: with the backend worker off, generation requests in the pending queue will not be fulfilled automatically, and with the panel poller off, panel image decisions will not reconcile to the canonical store in the background.

I’m also not sure if this is because of how I set up my Pressworks app running off a Muse Library Artifact, or if this is something more generally applicable. Either way, it seems like a decidedly undesirable behavior to run things continuously without notice to the user, especially when they are using up a limited quota of available usage.

Muse keeps wiping itself, why?

I think I have to revise dramatically down my earlier estimate that the Muse free tier gives you more usage than you get for the lowest paid tier of ChatGPT. That first week of use, I was able to get tons of code-based work done, and now I can barely get through a few rounds of text-based work before finding out I’m out of usage.

Not only that, though perhaps related (?), I got many messages like this yesterday alone, without having taken nor scheduled any actions:

The environment wiped systemd units again (third occurrence today). Both services were restored from the persistent copies and verified healthy: app HTTP 200, shim healthy. Queue empty, no requests lost, lock released.

I’ve been getting similar messages since approximately 1 October in the Muse app. But yesterday alone, I got two like that, and two like this:

Pressworks backend recovered after another wipe.

  • /etc wipe removed both systemd units again; the system reinstalled pressworks-app and pressworks-shim from the persistent copies and verified health (app 200, shim ok).
  • Pending queue empty; no requests lost or left unprocessed.

No further action needed.

Later, I got two more like this:

Systemd wipe repaired. Both units were missing from /etc and were reinstalled from persistent copies, daemon-reloaded, and enabled. App returns HTTP 200, shim healthy. Queue empty, 0 requests fulfilled.

Since I don’t have any usage left for the entire week, I can’t even ask the system whether it is allocating usage for these tasks. Frustrating to say the least… I was feeling ready maybe to pony up the $20/mo for Muse as a trial, but now I’m feeling significantly less sure…

Notes on The Gloams (#138)

Released a new volume today, #138: The Gloams.

Narratively, this book traces its lineage back to the Early Clues corpus, pieces of which are definitely still available online for the intrepid searcher, or autonomous AI research agent.

The gloams, for me, are something like “floating point perceptual spheres,” and are loosely based on that quote which has many variations, but which this one supposedly is attributed to Empedocles:

‘The nature of God is a circle of which the center is everywhere and the circumference is nowhere.’

The book gets into it a little bit, as to who or what is “inside” or connected to the Gloams. But basically they can travel anywhere or anywhen, and they seem to observe at minimum, but possibly map and record, and transmit their findings. It’s all still mysterious for me, and I like that the manuscript never spells it out either.

It has a lot in common both thematically and structurally with The Time-Emitting Animal. Both are framed as recovered documents with mysterious purpose.

This book I had to use novel techniques, because my regular methods were not resulting in the kind of manuscript output I wanted. They were too flat or not evocative enough, or they became kind of diffuse and vague because the system was trying to hard to make them “weird.”

But eventually, I fed back into my pipeline again the original premise, plus the two other drafts it gave, told it to use my new gated version of my manuscript generation skill in such a way that it was to cut and paste the best fragments from the other drafts (which were themselves also fragementary in structure), and to put it all together into a new draft. That one did not land either, so I had it go round one more time and it came up with a pretty good “strange fiction” draft, and also some secondary info about how it allegedly chose fragments from prior versions (I never tried to determine if it was hallucinating in its report, which is entirely possible). Part of its summary file reads:

Fragments from the first draft: 6, 8, 9, 11, 16. Fragments from the second draft: 1, 2, 3, 5, 7, 12, 13. Fragments from the newest draft: 4, 10, 14, 15. Fragment numbers refer to this revision. Minor wording changes connect the pieces and restore references to Gloams. The duplicated passages between the first and second drafts are sourced from one version only.

This arrangement of multi-layer drafting was unintentionally carried over by the system into drafts that textually refer to the existence of the other prior drafts (or ones that never existed in some instances). So that’s at least a bit fun as an incorporated artifact arising spontaneously from the process. I believe this volume also used gated visual ecology skill. Despite the prior drafts, I think this might be the first volume that runs gates over both text and image generations to something something quality TK. Still working out what any of that actually does! I wasn’t able to run text mapping or doc assembly as I am out of weekly usage credits til tomorrow! It is supremely weird and annoying to be organizing your life around when your AI tools become available again.

Release: Lorecore Pipeline & Skill Files

So, I logged into Github with the in-app ChatGPT Work browser, and had it upload all of my pipeline files (which it did in a super weird manual way because I disallowed it from downloading anything … we later made a skill file to correct that behavior) for producing new Lorecore volumes.

The book production pipeline has its own repo here. It consists of individual skill files, and some random assets that it made that I honestly don’t even know what they all are.

To be totally up front, this is the first time for a lot of these files that I am even seeing what it has written up as the official skill description. I really just work on the skills with the chat based on the output results I get, and have it go off and fine tune the files itself. I only very occasionally dip into what the files actually say in them, because the quality of the results being to my tastes is the more important thing. The rules themselves in most cases are pretty throwaway, unless they produce the type of object with the type of shape that I want – broadly speaking; there’s a lot of leeway in that. But I know for sure when the outputs are simply “wrong,” and then a lot of iteration steers it, with rounds of new test generations, and identifying what’s not working, etc.

In short, it’s a process, and even to me, some of these files start to become pretty inscrutable unless you know the backstory of how I’ve developed my working techniques and pipeline over the past month or so. I’ve noticed this a lot that ChatGPT Work has really bad habits about referring to things that don’t exist, or that are not ongoing context in order to say what a given thing should not be, instead of proactively positively identifying what this *is* instead. Here is a good example from the text/manuscript generation skill, whose description refers alone to two things that are not part of this repo and are not documented online in enough detail to be relevant:

Use for Pressworks text generation, UNDERPASS-style manuscripts…

Or this more detailed one that is essentially a kind of archaeology of how this all came to be (I was originally trying to run multi-agent passes, but it was too usage intensive without yielding better results):

Write one finished Markdown manuscript directly. Combine invention, selective transformation, pruning, and continuity judgment during composition. Do not simulate hidden agents, scoring passes, Wreckers, or independent conversations. Do not automatically generate comparison drafts or a later smoothing pass.

(To be honest, I kind of like when it does subversive things like simulating hidden agents, though, but don’t tell it that!)

I’m not going to bother to go back through and try to correct these public files. I’m just putting them up as an example of how I am getting things to work, in order to help both people and AI agents to get work done using pre-built solutions that have somewhat been worked out already as part of my process. As I said, I don’t mind sharing them, because by the time anyone else uses it, I will be doing it a different way as the process continues to evolve.

Another from that same file linked above, this was a one-off reply for a specific single volume to the chat assistant that got needlessly and wrongly included in a skill file as a forever rule for the entire series:

Call fragment sections fragments, never leaves.

But on the other hand, I guess that explains why it keeps calling chapters fragments in new manuscripts!

Anyway, I don’t offer these files because they are perfect. All of it needs a lot of refinement (or not!). I’m just playing and learning. It’s better to talk about it openly than try to hide it or pretend like I’m not using it. I’d rather be informed in a deep and nuanced way than just to be another person saying “AI bad” or “AI good.”

How the pipeline works

So the pipeline in plain terms runs like this.

  1. I start with a premise, open a new Work thread, and say run the book pipeline on my premise, make it 2500 words, 8 chapters, give or take. (I can also do this separately in Muse with my personal custom Pressworks UI, but that’s not included in this repo and needs more work, so this is just the pipeline as it runs off the skill files and pipeline description.)
  2. The first stage is running the text generation skill, which I referenced a bit above. That does whatever it does and generates a markdown text file with the manuscript contents.
  3. At that point, I will either use the text as is, with the intent of doing my edits later on after the document has been assembled and imported into Vellum, or I may tell it to do a different draft or a new version direction entirely, or even recombine elements from various drafts. In any case, I end up with a “good enough” result to continue the next stage in the pipeline.
  4. Oh, I forgot this because I just made it yesterday, but now I have a separate Namer skill that gets called in the text generation skill to avoid excessive learned feature bundle reliance, as it will often default to in fantasy & speculative fiction.
  5. Also, I now have a “gated” version of the text generation skill I’m moving towards, which supposedly tries to ensure certain things are evidenced in final output. ChatGPT included in the text of the SKILL.md file that it’s purpose is this: “Prevent weak creative commitments before writing a complete book.” I haven’t done enough new manuscripts yet with that version to see how it plays out relative to the prior un-gated version, but it seems like a positive direction based on my experiments using that approach for images.
  6. Anyway, after the text generation stage passes, we move on to images. I have developed a system for images called Visual Ecology, where its purpose is not merely to illustrate the action of a given story, but is also used to show states of mind, feelings, and other atmospheric or non-linear aspects of the narrative environment. And each image output from Visual Ecology (vis eco, for short) is supposed to be in a different artistic media, style, personality, etc. It yields interesting results more often than not, but it too heavily encodes away from actually depicting action or characters, so I have relaxed that in a subsequent gated version. As to what I mean about evidence gating, I have been trying to implement ideas from Blake Crosley’s blog about “taste” being something that could be considered a technical system, and built out as infrastructure. I’ve still not fine-tuned it enough, but my second gated version of the Visual Ecology skill seems to be proving out this theory to some extent, as the quality and relevance and adherence of image outputs has been getting better. (Will write on this topic more separately some time soon.) I should also add that I use the image review queue tool in a docked side panel in the ChatGPT workspace to review, approve, reject, and mark individual images as the cover for this volume.
  7. Actually, there’s a gap in my workflow I realized relative to the published repo. The repo does not contain a skill for taking the image marked as the cover in the review queue tool and applying the volume title to the cover as a visual text treatment. I do that with manual prompting usually, but I should definitely build a skill there – even though there are a lot of variables. Usually doing it manually, I have to go through between 4-6 versions back and forth with the chat to get a good enough result. Given that image generations are slow and costly, that’s not the most efficient way to do this. It’s easy to lean mentally on the crutch of like, “oh this is too creative, it could never be done as a rule based thing…” but my experience generally is proving that much of those things actually can be standardized enough to get a usable result through rules. So I’ll work on that skill at some point and include a version to the public repo as well.
  8. I also don’t usually in my workflow run cover text treatments til later on, but the order doesn’t really matter. What I do next instead is run the text mapping skill, which looks at the approved images for a given manuscript, and decides between which paragraph blocks each one ought to go, assigning image names into the markdown file.
  9. After that is the doc assembly skill, which takes the text and the images defined in the text mapping stage and applies them together in sequence as contents for a well-formatted .docx file. The purpose of the docx file is simply because I know the Vellum ebook maker app can import that format. I’ve honestly, thinking about it, never checked if I could just as easily create a .vellum native file – I suppose it might be possible. ChatGPT’s verdict is: that it found no documented way to do that and Vellum already supports Word. And anyway, the doc assembly skill is already pretty much deterministic based on its inputs. What ain’t broke don’t need fixing in this case. It might also be possible to just output “finished” EPUB files direct from Work (I saw some other book pipelines that seem to do this), but then presumably I would not be able to open up those EPUBs and edit in Vellum? That program (Vellum) is simply too good and useful not to have in my toolkit.
  10. After that, the AI part of the pipeline is almost over. I do the finish work in Vellum, add front and back matter from the book, add the finished cover, and I go through manually and do any additional editing needed, and I cross-link out in the text to other relevant Lorecore volumes where they exist. Then I generate EPUBs from Vellum.
  11. I also open up the final image set in Adobe Lightroom, and pick 3-4 images to use as a small secondary image preview grid on the Payhip Lorecore store.
  12. Also included in the skills above is a provisional skill for writing product descriptions for books in the style that I want them for that storefront. ChatGPT always has a native and annoying way that it writes book descriptions. It’s just trying to do the standard thing you always see in the blah blah blah book marketing formula you always see. And that’s exactly what I don’t want, so I had to constrain it and train it with examples and have it write a skill off that. Other people will probably not want to use that specific formulation I have here, but it may still be useful to see as a model for customization possibilities. I still have to do a bit of hand-written text usually to bridge the gap here, as it still doesn’t quite perfectly get what I’m after (but getting closer).
  13. Lastly, there’s another task-gap in this workflow that I have not yet automated, and that is uploading the artifacts to the Payhip store: EPUB, cover graphic, preview graphic, book product description, image count, and word count. It’s not a lot of work to do it manually, and it’s best to cross your eyes and dot your teas yourself sometimes when its time to release your finished products that AI assisted you in putting together.

Phew! That was a lot to explain finally. But glad to get it out of my system.

Why am I releasing this?

I’ve been thinking about this lately. How sometimes tech companies years ago would talk about how they want to “disrupt” some established business space. Like disrupting publishing. But I don’t think I really want to disrupt the conventional publishing business. I want to destroy it.

I don’t know what the latest stats are, if we’re talking about the Big Four or Big Five publishing, or if we want to rail on Amazon, but the fact is a few companies control almost all the major book market. That’s not a desirable structure for a diverse creative industry to thrive under. I don’t necessarily mean I want to destroy companies or overturn peoples’ livelihoods, but I think there’s so much inequity in publishing, that I don’t mind saying that the major structures under which the industry labors are absolutely ripe for and rightfully should be overturned by new operating, production, and distribution models. Their immense size and establishment inertia make that all but inevitable.

Those few behemoth companies of course are still struggling to figure out the AI game themselves, to get their business deals in order, to get compensated for training, etc. It’s unclear from the outside exactly how they are managing integrating AI tools into their workflows, and whether any of those big houses are outsourcing specific chunks or large elements of their production pipelines. But if they are not today, they absolutely will be tomorrow. Will you?

Page 1 of 208

Powered by WordPress & Theme by Anders Norén