Just to report back briefly on what I proposed in my last post, about using Muse as an orchestrator to drive a book production workflow, that employs both a ChatGPT Work session, and a regular Chat session…
Well, it worked. A proof of concept anyway. The rest is just picking apart the process, and tuning the component actions.
What happened, in simple terms, was that Muse ran a plan that I think ChatGPT originally wrote maybe? In it, Muse acted as orchestrator to review the existing corpus and an analytical layer I’ve been developing that names and describes some promoted “licensed” canonical elements from the story-universe, and identifies relationships between them. So applying the user-supplied premise, it dips into the corpus and the more structured entity/relationship set (“registry”), and its job is to prepare a research packet which can be sent to a regular ChatGPT chat (not “Work”) session in order to 1) generate a manuscript based on the premise that loosely fits into the canon of the world, with an allowable amount of invention, and then 2) creates a series of images to illustrate the text.
The way that happens is that I gave Muse my OpenAI login credentials. I felt very nervous about this before doing it. And not long after completing the task I deleted it from the “Secure Store” (which I have my doubts always about how “Secure” anything really can be in a time such as this). Looking back, it seems like there’s an option to do a one-time code by email, which I will probably do the next time I run an experiment such as this.
Here’s an excerpted section from the research packet/instructions Muse sent into the ChatGPT creative generation thread:
Symbol registry
No registered symbolic units or approved connections apply to this premise. The registry was searched for the premise’s themes (complaint, worship, noise, defector, bureaucracy, silence, UFR, Autogenos) and returned nothing applicable. Work from the premise and the excerpts above only.
Anyway then it opens its own browser, and you can watch while this is going on, stop it, or take over the browser as it runs. Or you can also steer it with text in the chat as you see what’s happening. And then it relays it back into the chat that Muse is having in its browser with ChatGPT. It’s circuitous and strange (compared to say just doing this yourself manually in a browser), but it seems to work as a base pipeline run. The rest is just refinement (image generation series step is one that will take some fixing, for example).
Here’s a clip from the generated text within the manuscript – the whole thing is still too short to use on its own, but there are some fun parts like this:
A complaint was an allowed form of speech when shaped correctly. It was not a cry cast into the streets. It was not a loud declaration before others. It was a restrained utterance placed within the channels established for such things. The defector knew that public voices were to remain low and contained, and therefore chose the formal path.
Then those generated assets get “downloaded” in the browser by Muse, and they appear in the chat UI in the Muse app, ready to proceed with the next step, which Muse handles, the mapping of where each image ought to go relative to the blocks of text (paragraphs).
Sidenote: you can also see to some extent if you open up ChatGPT while Muse is driving it elsewhere, the semi-real time updates to the Muse x ChatGPT conversation as it happens – though it was laggy and not always updating. This is area for big improvement for certain…
I think there ought to be some more standardized protocols of how agents logging in as you on other services ought to be identifying themselves (almost like how radio stations have to routinely state their call sign every x minutes). And if we’re going to enter like multi-player chat situations, where say I dive into the delegated real-time Muse x Chat to intervene in real time between the orchestrator agent (Muse) and the delegate agent (ChatGPT), there ought to be some clear way to establish who is speaking, and what ultimate authorities they have relative to one another.
I’d also like to see more fine-grained controls around, if I’m using an AI to drive another AI, to more clearly spell out: new threads only, no access to prior conversations, identify yourself as an agent, restrict other abilities to xyz… and then most importantly be able to test, verify, and enforce those limits. (Perhaps some of that is possible, but the amount of gaslighting agents do requires many strict checks.) Because otherwise, this is going to get bloody messy very fast as we enter a strange new world of agents-upon-agents-upon-agents ad infinitum… Lots more to say on that topic all by itself.
Anyway, process wise, then Muse-as-orchestrator sends back the text, image assets, and the mapping of where they should go sequentially into a new ChatGPT Work session, and the only thing that Work has to do is follow the mapping plan and assemble everything into a Vellum-optimized .docx file, which it does a great job with generally speaking, as I’ve run numerous other tests of that part of the pipeline prior to this.
This is a big deal because prior to this, I was running the entire pipeline from A–>Z in Work, and it would use up most or sometimes all of a 5 hr usage window on a Plus plan. (That’s a sentence that should not have to exist.)
I had Muse validate usage of itself and of Work before and after running the pipeline with this orchestration. And I did not rigorously validate this on my own, or carefully track tokens or anything – so take this with a tremendous grain of salt. But Muse claimed that it used 4% of its weekly usage limit on a free plan, and that Work only used 1% of its weekly usage. Again, that will take some more rigorous measurement at some point (and then optimization for efficiency), but for now, I just wanted to get everything working, and I have succeeded in that, and it’s really interesting!
Leave a Reply
You must be logged in to post a comment.