Getting the First 69% For Free With Decision Models

I zero-shot labeled much of a bioreactor dataset using an open-weight decision model, local compute on my laptop, and prompts so basic that they're a little silly.

Share
A big pile of blocks that say "WHAT?". Here we're trying to ask what a datastream represents. It's clever.
Photo by Vadim Bogulov / Unsplash

When you pull data into an automated system, the first task is often to ask "What is this?" Decision models can be a fast, resource-light way to automate answering this question in standardized way. In this proof-of-concept, I zero-shot labeled 69% of fields correctly from an AMBR250 bioreactor dataset using the open-weight Laya model, local compute on my laptop, and prompts so basic that they're a little silly.

The model was trivial to implement and performed very well compared against other trivial approaches

My big takeaway from this work is that these decision models are already easy to get running, resource-efficient, and good enough to be useful. They are currently limited by an inability to reason about numerical data and can struggle with niche biotech terms, but those issues can be overcome. These models should be considered seriously in any automation data pipeline, especially when speed is crucial or you need to run locally.

To achieve high reliability with these models, the juicy targets are:

  • Improve the ontology that we use to describe the bioreactor's world
  • Fine-tune the model for biotech language
  • Integrate numerical models that infer more context from the actual data

Read on for my exploration of how to use these models effectively and a case study using ABPDU's AMBR250 data. Calculations are available on GitHub.

Premise

I've been seeing a lot of discussion of "System 1" decision models since Jev from TypeSafe launched. OpenAI also just announced an API following this pattern. There's a lot of writing about these models and, from what I can tell, much of it is marketing hype. The important thing for our purposes is that these models are designed to output one of a constrained set of choices rather than arbitrary text and, as a result, require orders of magnitude fewer resources. You can even run them on a laptop CPU and get workable performance.

This use case sounds an awful lot like a specific need I see when handling biotech data for automated consumption - data comes in with arbitrary labels on it that humans (hopefully) understand, and we need to apply canonical labels that our automated system knows about and can act on. This is a crucial step whether we're planning to store the data for later analysis, pass it to a model to reason about, or feed it directly back into a self-driving lab to take actions.

Everyone has to solve this problem eventually (coincidentally Invert Bio recently posted about it!) because at some point we all run into a giant pile of spreadsheets with thousands of names like Pump B pulse flow on time_SP (s) and we have to figure out what they all mean.

Aw man.

Since this technology seems to match this application pretty well, I set out to see what kind of performance I could get from it.

Hypothesis

My initial question for this work was "Can I use a Decision Model to canonicalize a data stream based on just the data, not looking at the labels at all?"

As a domain expert, I know that there are patterns in data that help me pick out what it likely represents - Temperature is usually in biologically-favorable ranges, Stir Rate tends to trend up over a run, etc. Since these models use similar architectures as LLMs, which tokenize linear text and extract that information into concepts, maybe they could work with linear data as well?

To spoil the outcome, the answer was no - it turns out Laya is very bad at working with numbers. BUT, it proved surprisingly good at inferring information about a dataset just by looking at a small amount of nonstandard metadata. Providing units of measure was often enough for the model to make a good guess at what the data represented.

The model was basically useless unless there was a raw label or metadata like units included

So I asked a follow-up question: "Can I use an off-the-shelf Decision Model to save a meaningful amount of effort in canonicalizing a raw labeled dataset?"

Specifically I wanted to use an existing model without retraining it. I'm trying to test the current state of the art for practitioners of biotechnology, not machine learning engineers. So we're focused on what anyone reading this could get running today if they wanted to.

Approach

In line with the goal that you could run this today, I chose to use an open-weight model that you can download from Hugging Face right now without putting in a credit card. These models also have the advantage that none of your data needs to exit your network if you run locally. Speed, cost, and cybersecurity are uniquely important in manufacturing environments, so these models would be attractive on all fronts if they could get good results. I chose the Laya model because it's open-weight and I had heard of it.

Staying on the theme of practical use, I took an unscientific approach. The goal here was to build intuition around the value of the model, not benchmark rigorously.

Setting Up The Model With Mock Data

My approach was to make up some plausible-seeming datasets that I might recognize as a domain expert, then ask in plain language "What type of measurement does this appear to be?". I described 4 possibilities for the model to choose from, pH, Temperature, Stirrate, and Other. I tried to use natural language descriptions focused on the data. Below is my description of Temperature:

The temperature of the broth. Represented as a list of numbers that are usually between 20 and 40, but sometimes goes out of that range. Often named 'T', 'Temp', or some variation of that

I then ran the inference on the model, which is just a couple of lines of Python:

from laya import Router

router = Router()

state = '{"name": [30.0,30.9,31.5,29.8]}'
questions = my_questions_json #JSON with descriptions like the one above

result = router.predict(state, questions)

The output of the prediction function is a JSON object with a choice, a confidence level for each option, and several other pieces of metadata. Here's the output of my pretty-print function for the code above:

The model's ability to report confidence is a crucial feature, because it lets us decide if we should act on the model's choice or if we should pass the data on to some other system for further review. In my analysis, I picked an arbitrary confidence gate of 0.5 - if the model reported a confidence ≤0.5, I marked it as "needs review." In practice we might ask a human to review it, or we might have a secondary system like an LLM do a more rigorous assessment before trying again.

Testing Against Real Data

After getting a sense of what I could expect from the model, I tried to canonicalize ABPDU's public Rhodosporidium toruloides dataset. This is a data dump of a set of AMBR250 runs where each datastream is represented by a CSV file with a single line describing the data, like Stir average power (W).

The AMBR250 system reports its data in great detail. To describe all the data it provides, I needed a more complicated model of measurements than a single label. I broke every measurement down into:

  • Name: The canonical name of the measurement, e.g. "Stirrate"
  • Submeasurement: The value of interest relating to the name, e.g. "Power"
  • Unit of Measurement: The units the value is expressed in, e.g. "Watts"
  • Process Location: Where in the process the value is referring to, e.g. "Fermentation Broth"
  • Signal Type: Is this a Process Variable (something measured), a Setpoint (a desired value), or a Command Output (a signal coming from the controller?

Learning from the previous work, I used very simple prompts. For each name option, I used the following description:

Related to [name] in a bioreactor

I ran the inference on all 340 measurements reported for a single reactor using an adapted version of the code above, then compared the results to a hand-labeled "correct" dataset.

Outcome

The model did surprisingly well considering how little work it was to implement. Counting each of the 5 questions as a label, it correctly labeled 69%, confidently mislabeled 6%, and reported low confidence on 25% of them. In a real workflow, we would pass the low-confidence records to a secondary system, but the confidently-mislabeled records are a concern.

Ultimately, the Decision Model can save meaningful time labeling, but off-the-shelf it can't be used as a complete solution

This is a vast improvement over methods requiring similar effort and resources: the best alternative I could imagine that takes just a few lines of code is to always guess the most common value - shown here as "Best Fixed Guess". So for name we might always enter "Flowrate", since the data is skewed to have more of those. Obviously this is going to get a lot wrong, but it really is only slightly more easier to implement than the decision model approach.

A more realistic comparison would be to approaches that may perform better, but use more resources - either comparing to the tokens/cost of an LLM agent or the developer resources required to make a deterministic classifier system. Neither is data we have, but I feel confident that the decision model approach wins on resource demand by several orders of magnitude and loses on accuracy by some less drastic amount in both cases.

The question, then, is: what effort would be required to upgrade this low-resource/middling-accuracy approach to a low-resource/high-accuracy approach?

Unit Labeling Isn't Acceptable

% of records the units classifier labeled with each unit of measure. It was much too eager to label entries as milliliter

If you're carefully following the math, you'll see that I actually completely discounted the model's classification of units of measure entirely. In this experiment it labeled almost everything as milliliter, and applied other labels seemingly randomly. The formulation of the prompt can strongly affect the prediction, and I found that when the model isn't well-tuned for the question you're asking, the prompt can cause it to gravitate towards one answer that happens to be a statistically "good" generic answer.

I think it's possible that a well-tuned prompt could get good answers out of it, but units are a relatively constrained space, so I think it might be more fruitful to use a purpose-built tool (Have you tried Pint? It's pretty good!)

Ontology Matters

Because we're using the model for classification, it can only be as good as the classes provided to it. I made up my classes ad hoc based on a mix of what I know is common in bioreactor data and what I saw was in the data. Maybe more impactful, I made up the class categories - the ontology - just as haphazardly. For many of the records, I found that I couldn't define how they fit into my ad hoc ontology, even if I did conceptually understand what the label meant.

We see that for name, submeasurement, and process_location , a little more than 10% of records simply had no good defined answer. This meant that nearly 30% of records had no hope of getting all categories labeled perfectly regardless of the model's performance.

These numbers don't capture further problems we might see if, for instance, name and submeasurement are the wrong categories in the first place. Maybe it should be instrument and metric. Developing the ontology is a balance between modeling what's available and what's important in the application. It typically takes a lot of time. If we had a good baseline ontology for bioreactors to start from, it could impact performance meaningfully (Have you tried PREFER? I haven't yet. Maybe it's good!).

Performance Improved When Handling Confidence and Ontology

The data supports the idea that, before touching anything about the model, we can get better practical performance out of this off-the-shelf version by handling low-confidence decisions and those that don't fit into our ontology.

Every question greatly outperformed random chance, but overall the accuracy is not high enough to solve the labeling problem entirely. This can be partially mitigated by only taking confident rows or improving the ontology to get more of the rows definable.

This chart shows that every category improves meaningfully if you only take the entries where the model is confident ("confident_rows") and where the ontology provides a correct answer ("defined_rows").

The defined rows were arrived at by excluding from the calculation any record for which I couldn't find a correct answer when manually labeling. As discussed above, the possible improvement achievable by actually improving the ontology is on the spectrum from "much better" to "slightly worse".

If we assume we have a good fallback system for low-confidence entries, then we already get major improvements by taking only high-confidence entries. I chose a confidence gate arbitrarily at 0.5. A broader study finding the optimal confidence gate that balances bad labels and unknowns would likely improve outcomes even further.

Both confidence and ontology improvements have multiplicative impacts that are apparent in the correct_overall column - these are scores for when an entry gets all categories correctly labeled. Since one bad apple can spoil the bunch, smarter handling for confidence and ontology edge cases could make a big difference in the number of full records that are zero-shot correct.

General Decision Model Usage

In general, I found the model worked best on metadata that included verbal, conceptual information, especially when the concept was something that might have been in the model's training dataset - so Temperature (°C) worked well, but Dissolved Oxygen (% air sat.) was probably too niche.

Nothing I saw suggests the model has any way to reason about numbers in the dataset. In fact, literally sending random ASCII characters to the model sometimes seemed to get me similar results to sending numerical data.

The specific working of the prompt strongly affects outcome in this base model. I found the best thing to do was rely on existing concepts rather than prompt-engineer for hours. I suspect fine-tuning the model on biotech data would get more consistent results.

Takeaways

  • Decision Models represent a low-cost/middling-accuracy approach that makes them worth considering as part of an automated biotech data pipeline.
  • They likely add value if you're hand-cleaning data and want to knock out the low-hanging fruit without doing a lot of setup. You'll save time even after manually correcting.
  • For an automated system where it's impractical to do human review, the output would need to be passed to a downstream model for validation and correction - could be an LLM, a numerical tool that uses heuristics, or just a set of deterministic rules.
  • Your mileage may vary: Numerical rules basically don't work because it's not what the model was built for. Stick to simple prompts if you don't plan to fine-tune the model.
  • This approach is easy to get started and cheap to maintain - it can run reasonably well on a regular old laptop, and would run fast and cheap on a cloud server. This alone makes it worth trying.

Further Questions

  • Would an improved ontology get better results? By how much?
  • How much improvement could we get by fine-tuning the model?
  • What numerical approaches could we pair with this model, since it isn't designed to work with numbers?
  • Is there a way to prompt that would get better results? How can this be studied systematically?
  • Would alternative decision models perform better? Is it worth the tradeoffs of using a cloud service like Jev?

Detail and Code on GitHub

All of my work is available in a Jupyter notebook on GitHub. Feel free to play with your own data and adapt it to your needs.