Mapping what’s next for conversational AI in Alchemite™

This is the last of a series of three ‘blogs from the road’ from Joel Strickland, as he worked remotely while travelling through South and Central America. As he completed is own adventurous mapping exercise, he considered what a map of the future for conversational AI in the Alchemite™ machine learning platform might look like.

Blog post by Joel Strickland.

I’m writing this in Rurrenabaque, on the edge of a few days offline in the Bolivian Amazon. Before I disappear for a bit, I wanted to write down one idea that has become central to how we think about Alchemite’s next step.

Over the past few years, we’ve been looking closely at what people actually want from agentic AI in scientific software. We’ve now written up that work in a preprint: Talk Freely, Execute Strictly: Schema-Gated Agentic AI for Flexible and Reproducible Scientific Workflows.

The paper is broader than Alchemite, but its central question is directly relevant to where we think Alchemite should go next:

How do you make scientific software more flexible and conversational without making it less reproducible?

That now looks like the core challenge.

In our interviews, practitioners consistently wanted natural-language interaction to express intent, refine analyses, find the right workflow, and reduce the friction of rigid interfaces.

But they wanted something else just as strongly: determinism.

Anything contributing to the scientific record still needs to be stable, repeatable, and grounded in well-defined operations. People may want to ask questions naturally and change direction halfway through, but once the system is doing real work, they still need to know what happened, why it happened, and whether they could reproduce it later.

In the preprint, we describe this as a tension between conversational flexibility and execution determinism. Figure 1 shows that pattern clearly in the interview data. What people were asking for was not just more conversational software. They were asking for software that is easier to work with, without losing the control, traceability, and reproducibility that matter once the work becomes scientifically meaningful.

Figure 1. Interview responses clustered around two dominant requirements: conversational flexibility and execution determinism.

That is what makes this more than a UX question.

A scientist might want to describe a goal in natural language, but still expect the resulting workflow to run through validated, reproducible steps. So the challenge is not just making the interface feel more natural. It is designing the system so that flexibility at the point of interaction does not come at the cost of reliability underneath.

Figure 2 captures that trade-off more directly. If you optimise too far for conversational flexibility, you can end up with systems that feel fluid and capable, but are harder to validate in advance. If you optimise too far for execution determinism, you get systems that are robust and auditable, but rigid to use.

That is the design space we think matters. The challenge is not choosing one side or the other. It is finding ways to increase flexibility without giving up determinism.

Figure 2. Conversational flexibility and execution determinism define the core architectural trade-off.

That is exactly why this matters for Alchemite.

Alchemite already starts from the determinism side of the equation, and that is a large part of its value. It gives users robust, validated, reproducible analytical workflows, and that foundation is not something we want to loosen.

What we do want is to add flexibility on top of that foundation.

Not flexibility in the sense of letting an open-ended model improvise its way through scientific work, but flexibility in the sense of making workflows easier to find, navigate, and use. The goal is to make it simpler for users to express intent in natural language and get to the right analysis more quickly, while keeping execution grounded in the same standards of rigour.

People are not really asking scientific software to become more like a generic assistant. They are asking it to become easier to work with, while remaining just as trustworthy.

In a general assistant, variation can be part of the appeal. In scientific software, it is usually a liability. Once the system is doing real work, trust depends on knowing what happened, why it happened, and whether it can be repeated. You do not get that by hoping the model behaves itself. You get it from system design.

So for us, the opportunity in Alchemite is not to replace rigorous workflows with open-ended agents. It is to extend rigorous workflows with a more natural interface.

If we get that right, conversational flexibility should lower onboarding effort, help users get more out of the platform, and make it easier to move from question to answer without having to learn every layer of the interface up front. That should mean faster adoption, less friction in day-to-day use, and more value from workflows users already trust. But those benefits only matter if the execution underneath remains dependable.

To me, that is the interesting direction for agentic AI in scientific software.

Not improvisation.

Not autonomy for its own sake.

And not a chatbot pasted onto a serious platform.

Something better than that: software that lets people work more naturally while keeping the parts that matter most stable and trustworthy. That is the balance we want to bring into Alchemite.

Reference

Strickland, J., Vijeta, A., Moores, C., Bodek, O., Nenchev, B., Whitehead, T., Phillips, C., Tassenberg, K., Conduit, G. and Pellegrini, B., 2026. Talk Freely, Execute Strictly: Schema-Gated Agentic AI for Flexible and Reproducible Scientific Workflows. arXiv preprint arXiv:2603.06394.

Search