So, I logged into Github with the in-app ChatGPT Work browser, and had it upload all of my pipeline files (which it did in a super weird manual way because I disallowed it from downloading anything … we later made a skill file to correct that behavior) for producing new Lorecore volumes.

The book production pipeline has its own repo here. It consists of individual skill files, and some random assets that it made that I honestly don’t even know what they all are.

To be totally up front, this is the first time for a lot of these files that I am even seeing what it has written up as the official skill description. I really just work on the skills with the chat based on the output results I get, and have it go off and fine tune the files itself. I only very occasionally dip into what the files actually say in them, because the quality of the results being to my tastes is the more important thing. The rules themselves in most cases are pretty throwaway, unless they produce the type of object with the type of shape that I want – broadly speaking; there’s a lot of leeway in that. But I know for sure when the outputs are simply “wrong,” and then a lot of iteration steers it, with rounds of new test generations, and identifying what’s not working, etc.

In short, it’s a process, and even to me, some of these files start to become pretty inscrutable unless you know the backstory of how I’ve developed my working techniques and pipeline over the past month or so. I’ve noticed this a lot that ChatGPT Work has really bad habits about referring to things that don’t exist, or that are not ongoing context in order to say what a given thing should not be, instead of proactively positively identifying what this *is* instead. Here is a good example from the text/manuscript generation skill, whose description refers alone to two things that are not part of this repo and are not documented online in enough detail to be relevant:

Use for Pressworks text generation, UNDERPASS-style manuscripts…

Or this more detailed one that is essentially a kind of archaeology of how this all came to be (I was originally trying to run multi-agent passes, but it was too usage intensive without yielding better results):

Write one finished Markdown manuscript directly. Combine invention, selective transformation, pruning, and continuity judgment during composition. Do not simulate hidden agents, scoring passes, Wreckers, or independent conversations. Do not automatically generate comparison drafts or a later smoothing pass.

(To be honest, I kind of like when it does subversive things like simulating hidden agents, though, but don’t tell it that!)

I’m not going to bother to go back through and try to correct these public files. I’m just putting them up as an example of how I am getting things to work, in order to help both people and AI agents to get work done using pre-built solutions that have somewhat been worked out already as part of my process. As I said, I don’t mind sharing them, because by the time anyone else uses it, I will be doing it a different way as the process continues to evolve.

Another from that same file linked above, this was a one-off reply for a specific single volume to the chat assistant that got needlessly and wrongly included in a skill file as a forever rule for the entire series:

Call fragment sections fragments, never leaves.

But on the other hand, I guess that explains why it keeps calling chapters fragments in new manuscripts!

Anyway, I don’t offer these files because they are perfect. All of it needs a lot of refinement (or not!). I’m just playing and learning. It’s better to talk about it openly than try to hide it or pretend like I’m not using it. I’d rather be informed in a deep and nuanced way than just to be another person saying “AI bad” or “AI good.”

How the pipeline works

So the pipeline in plain terms runs like this.

  1. I start with a premise, open a new Work thread, and say run the book pipeline on my premise, make it 2500 words, 8 chapters, give or take. (I can also do this separately in Muse with my personal custom Pressworks UI, but that’s not included in this repo and needs more work, so this is just the pipeline as it runs off the skill files and pipeline description.)
  2. The first stage is running the text generation skill, which I referenced a bit above. That does whatever it does and generates a markdown text file with the manuscript contents.
  3. At that point, I will either use the text as is, with the intent of doing my edits later on after the document has been assembled and imported into Vellum, or I may tell it to do a different draft or a new version direction entirely, or even recombine elements from various drafts. In any case, I end up with a “good enough” result to continue the next stage in the pipeline.
  4. Oh, I forgot this because I just made it yesterday, but now I have a separate Namer skill that gets called in the text generation skill to avoid excessive learned feature bundle reliance, as it will often default to in fantasy & speculative fiction.
  5. Also, I now have a “gated” version of the text generation skill I’m moving towards, which supposedly tries to ensure certain things are evidenced in final output. ChatGPT included in the text of the SKILL.md file that it’s purpose is this: “Prevent weak creative commitments before writing a complete book.” I haven’t done enough new manuscripts yet with that version to see how it plays out relative to the prior un-gated version, but it seems like a positive direction based on my experiments using that approach for images.
  6. Anyway, after the text generation stage passes, we move on to images. I have developed a system for images called Visual Ecology, where its purpose is not merely to illustrate the action of a given story, but is also used to show states of mind, feelings, and other atmospheric or non-linear aspects of the narrative environment. And each image output from Visual Ecology (vis eco, for short) is supposed to be in a different artistic media, style, personality, etc. It yields interesting results more often than not, but it too heavily encodes away from actually depicting action or characters, so I have relaxed that in a subsequent gated version. As to what I mean about evidence gating, I have been trying to implement ideas from Blake Crosley’s blog about “taste” being something that could be considered a technical system, and built out as infrastructure. I’ve still not fine-tuned it enough, but my second gated version of the Visual Ecology skill seems to be proving out this theory to some extent, as the quality and relevance and adherence of image outputs has been getting better. (Will write on this topic more separately some time soon.) I should also add that I use the image review queue tool in a docked side panel in the ChatGPT workspace to review, approve, reject, and mark individual images as the cover for this volume.
  7. Actually, there’s a gap in my workflow I realized relative to the published repo. The repo does not contain a skill for taking the image marked as the cover in the review queue tool and applying the volume title to the cover as a visual text treatment. I do that with manual prompting usually, but I should definitely build a skill there – even though there are a lot of variables. Usually doing it manually, I have to go through between 4-6 versions back and forth with the chat to get a good enough result. Given that image generations are slow and costly, that’s not the most efficient way to do this. It’s easy to lean mentally on the crutch of like, “oh this is too creative, it could never be done as a rule based thing…” but my experience generally is proving that much of those things actually can be standardized enough to get a usable result through rules. So I’ll work on that skill at some point and include a version to the public repo as well.
  8. I also don’t usually in my workflow run cover text treatments til later on, but the order doesn’t really matter. What I do next instead is run the text mapping skill, which looks at the approved images for a given manuscript, and decides between which paragraph blocks each one ought to go, assigning image names into the markdown file.
  9. After that is the doc assembly skill, which takes the text and the images defined in the text mapping stage and applies them together in sequence as contents for a well-formatted .docx file. The purpose of the docx file is simply because I know the Vellum ebook maker app can import that format. I’ve honestly, thinking about it, never checked if I could just as easily create a .vellum native file – I suppose it might be possible. ChatGPT’s verdict is: that it found no documented way to do that and Vellum already supports Word. And anyway, the doc assembly skill is already pretty much deterministic based on its inputs. What ain’t broke don’t need fixing in this case. It might also be possible to just output “finished” EPUB files direct from Work (I saw some other book pipelines that seem to do this), but then presumably I would not be able to open up those EPUBs and edit in Vellum? That program (Vellum) is simply too good and useful not to have in my toolkit.
  10. After that, the AI part of the pipeline is almost over. I do the finish work in Vellum, add front and back matter from the book, add the finished cover, and I go through manually and do any additional editing needed, and I cross-link out in the text to other relevant Lorecore volumes where they exist. Then I generate EPUBs from Vellum.
  11. I also open up the final image set in Adobe Lightroom, and pick 3-4 images to use as a small secondary image preview grid on the Payhip Lorecore store.
  12. Also included in the skills above is a provisional skill for writing product descriptions for books in the style that I want them for that storefront. ChatGPT always has a native and annoying way that it writes book descriptions. It’s just trying to do the standard thing you always see in the blah blah blah book marketing formula you always see. And that’s exactly what I don’t want, so I had to constrain it and train it with examples and have it write a skill off that. Other people will probably not want to use that specific formulation I have here, but it may still be useful to see as a model for customization possibilities. I still have to do a bit of hand-written text usually to bridge the gap here, as it still doesn’t quite perfectly get what I’m after (but getting closer).
  13. Lastly, there’s another task-gap in this workflow that I have not yet automated, and that is uploading the artifacts to the Payhip store: EPUB, cover graphic, preview graphic, book product description, image count, and word count. It’s not a lot of work to do it manually, and it’s best to cross your eyes and dot your teas yourself sometimes when its time to release your finished products that AI assisted you in putting together.

Phew! That was a lot to explain finally. But glad to get it out of my system.

Why am I releasing this?

I’ve been thinking about this lately. How sometimes tech companies years ago would talk about how they want to “disrupt” some established business space. Like disrupting publishing. But I don’t think I really want to disrupt the conventional publishing business. I want to destroy it.

I don’t know what the latest stats are, if we’re talking about the Big Four or Big Five publishing, or if we want to rail on Amazon, but the fact is a few companies control almost all the major book market. That’s not a desirable structure for a diverse creative industry to thrive under. I don’t necessarily mean I want to destroy companies or overturn peoples’ livelihoods, but I think there’s so much inequity in publishing, that I don’t mind saying that the major structures under which the industry labors are absolutely ripe for and rightfully should be overturned by new operating, production, and distribution models. Their immense size and establishment inertia make that all but inevitable.

Those few behemoth companies of course are still struggling to figure out the AI game themselves, to get their business deals in order, to get compensated for training, etc. It’s unclear from the outside exactly how they are managing integrating AI tools into their workflows, and whether any of those big houses are outsourcing specific chunks or large elements of their production pipelines. But if they are not today, they absolutely will be tomorrow. Will you?