One of the main things I’ve been struggling with as my comprehension and ability to execute on my ideas with AI has grown is: when to use each element?
Right now, I guess I see five different distinct ways to solve the same sets of task-based or process/workflow-based problems. They go something like this:
- Use natural language prompting in sequence in order to achieve a specific goal. Honestly, in many cases, this still works just fine (to the extent that it works at all, which is maybe debatable separately), especially for unclear tasks or one-off goals. It’s a good way to explore and see what can work. And anyway, it’s the basis of all the other methods. But, critically, if you’re trying to set up processes that you can run repeatedly, and get more or less consistent results, then you might at worst run into complications, or at best, spend a lot more time for the same quality of outputs. Whether that’s an issue depends of course on the specific use case.
- Develop a skill file to re-use for certain repeated tasks with clearly defined parameters as to process and outcome. I guess I partly answered the “when to use” question for that approach above. But it’s still not entirely obvious to me, which is why I wanted to pick this all apart. I guess I would say that if the task is not a one-off task, but one you’ll run again and again, then setting up a skill makes a certain amount of sense. If you don’t know how to make skills in your workspace, my experience has been, you just prompt the assistant with something like “Make a skill that does x with abc elements…” and then the assistant in a compatible system will just do it for you. But then you will need to run the skill on real example work-pieces a few times, identify what is not working, and have the system tweak it.
- Put together a pipeline that runs a series of skills in sequence. Again, I guess I’m answering my own question with that as a header, but my pipelines typically consist of a number of different skills which, when taken together, and with approval checkpoints built in, end with a complex piece of work that gets produced of an adequate baseline quality. It might not be the finished product, and it may take some handwork still in my case, as I’m not merely trying to just repeat a formula, but it gets me most of the way there. And this is where testing and fine tuning each of the component skills that you are chaining together becomes even more important. If you have a subsequent step that relies on appropriate outputs from a previous step, but you have not nailed down perfectly the skill that runs the previous step, then you’re likely to pollute the rest of your pipeline and not achieve the desired result.
- Develop a persistent UI control surface that visualizes (dashboard) and/or controls elements of a prompt-based workflow or more formalized skill-based pipeline, and enables you to more easily track and modify state over time. For me, this boiled down to versioning of manuscripts, and the need to track approvals and rejections for both text and for image rounds. Doing this purely in a chat-based flow was becoming a nightmare of scrolling back and forth and fighting with the system about which ones I had already approved and should be included in output stages. What’s interesting is that once you do develop persistent control surfaces, you can both use them as a UI itself, but also interact with it and modify the UI through chat-based turns that happen alongside.
- Delegate an agent to handle the whole thing for you (or components). This was my experiment around using Muse to drive ChatGPT to split up elements of my pipeline work between the two of them, based on what stage each was good at (or in which system I had done the most prior work to support a given stage). I had encoded my skills and full pipeline already in ChatGPT Work first, though I replicated and tinkered with them in Muse also to compare, and determine who was better at each element. Then Muse acted as an orchestrator to run a sort of meta-pipeline that assigned different parts to each system, and pulled the results of each stage back into Muse, before sending them onto the next leg of the journey. Now, I guess there are two major ways to approach delegating complex tasks to an agent: one is the skill/pipeline based method I’ve described above, which I consider a more structured method. The other would be unstructured (or maybe semi-structured), where you define the end goal (a book of 2000 words on a given topic, split into 8 chapters, with one image each per chapter, delivered as a .docx file), and then the assistant/agent determines the steps to take on its own. And then you could perhaps have the agent write from that a more durable set of skills fit into a sequential pipeline to re-use. This might work too.
The thing is, I guess, all of these things *might work.* And that’s kind of the problem. But maybe also part of the fun of the whole thing too: it’s completely open-ended, and many paths might lead to similar results. And the most “correct” approach is probably some combination of all the above.
Leave a Reply
You must be logged in to post a comment.