Databricks ai_decide Explained: Fast AI Decisions in SQL

Not every AI task needs a model that writes. Here is how Databricks' decision function routes, tags and scores your data in plain SQL.

Which team should handle this ticket? Does this document need a human to look at it? Should this prompt go to a small model or a big one?

These are small questions. But in a real company, they get asked millions of times a day.

For the last few years, most teams answered them the same way. They sent the text to a large language model and asked it to write an answer. Then they wrote extra code to pull one word out of that answer.

It works. It is also slow and expensive for what is really a simple choice.

Databricks now has a function built for exactly this job. It is called ai_decide.

https://youtu.be/9YwY-2GNTnc

The problem with using a writer to make a choice

A large language model is a writer. It produces text one piece at a time.

When you ask it "Is this ticket about billing or shipping?", it does not just pick. It writes. Sometimes it writes "Billing." Sometimes it writes "This appears to be a billing issue." Sometimes it adds a polite sentence at the end.

So the data engineer adds a cleanup step. Strip the spaces. Lowercase the text. Check the value is one of the allowed labels. Retry when it is not.

Now multiply that by ten million rows. You pay for every word the model writes. You wait for every word too. And you still get the odd answer your parser did not expect.

The question was never "please write something". The question was "pick one". That is a different kind of work.

What ai_decide is

ai_decide is a built-in Databricks AI Function. You call it from SQL, just like UPPER() or COALESCE().

It is powered by a decision model rather than a text generation model. Databricks describes it as taking unstructured input and a set of questions, and returning decisions and their probabilities directly, in a fraction of a second (Databricks blog).

That changes three things.

It is fast. There is no long answer to write, so there is no long answer to wait for.

It is structured. You get back a clean result with known fields. No parsing of free text.

It gives you numbers. Each answer comes with probabilities, so you can see how sure the model is, not just what it picked.

At the time of writing, ai_decide is in Beta. Workspace admins turn it on from the Previews page (Databricks docs).

The shape of a call

The syntax is short.

ai_decide(state, questions [, options])

state is the thing you want to judge. It can be plain text, a JSON string, or a VARIANT that came out of another AI function such as ai_parse_document or ai_extract. It can be different on every row.

questions is a JSON object. Each key is a question name you choose. Each value says what kind of question it is and how to answer it. The same questions are applied to every row.

options is optional. Today the only option is the version, which defaults to '1.0'.

One call can ask several questions at once about the same input. That is a nice detail. You do not need three separate model calls to get a category, a yes or no, and a priority.

Three kinds of questions

This is the heart of the function. Every question is one of three types.

1. noul: how likely is this true?

A noul question returns a probability between 0 and 1.

Think of it as a yes or no question where you get to see the confidence.

{
  "needs_escalation": {
    "type": "noul",
    "instructions": "Does this ticket need immediate escalation?",
    "criteria": {
      "true": "An active service outage blocks the customer from working.",
      "false": "The customer can continue working or has a workaround."
    }
  }
}

The criteria part is optional. But describing what "true" and "false" mean in your business is where most of the quality comes from. "Urgent" means different things in different companies.

The answer looks like this:

{ "needs_escalation": { "type": "noul", "probability": 0.1 } }

2. choice: pick one label

A choice question picks one label from a list you define. You can have between 1 and 255 labels.

{
  "team": {
    "type": "choice",
    "instructions": "Which team should handle this ticket?",
    "criteria": {
      "shipping": "Delivery or shipment issues",
      "billing": "Payment or invoice issues",
      "technical_support": "Product technical problems"
    }
  }
}

The answer gives you the winner, the probability for every label, and a confidence score.

{
  "team": {
    "type": "choice",
    "choice": "shipping",
    "probabilities": { "shipping": 0.8, "billing": 0.1, "technical_support": 0.1 },
    "confidence": 0.9
  }
}

The probabilities always add up to 1. The label is always one of yours, spelled exactly as you wrote it. That alone removes a whole class of cleanup code.

3. score: place it on a scale

A score question uses an ordered list of 2 to 10 descriptions, from lowest to highest.

{
  "urgency": {
    "type": "score",
    "instructions": "How urgent is this ticket?",
    "criteria": [
      "Routine request with no time pressure",
      "Time-sensitive issue with a workaround",
      "Critical issue that blocks the customer"
    ]
  }
}

The result is a weighted average of the positions. The first description is 0, the next is 1, and so on.

So if the model gives 10% to "routine", 30% to "time-sensitive" and 60% to "critical", the score is 1.5. That number can be fractional, which is useful. A score of 1.9 and a score of 1.1 are both "level 1" if you round, but they are not the same ticket.

A full example

Here is the example from the Databricks documentation. It asks three different questions about one product listing in a single call.

SELECT
  ai_decide(
    '{"name": "TrailShell jacket",
      "description": "Lightweight waterproof hiking jacket made from recycled polyester. Packs into its own pocket."}',
    '{
      "category": {
        "type": "choice",
        "instructions": "Which product category best fits this item?",
        "criteria": {
          "outerwear": "Jackets, coats, and other protective outer layers",
          "footwear": "Shoes, boots, and sandals",
          "accessories": "Bags, hats, and other accessories"
        }
      },
      "recycled_materials": {
        "type": "noul",
        "instructions": "Does the listing state that the product uses recycled materials?"
      },
      "hiking_suitability": {
        "type": "score",
        "instructions": "How suitable is this product for hiking in rainy weather?",
        "criteria": [
          "Not suitable for outdoor use in rain",
          "Offers some protection from rain",
          "Designed for hiking with waterproof protection"
        ]
      }
    }',
    map('version', '1.0')
  ) AS decision;

The result comes back as a VARIANT with three fields: response, metadata and error_message. On success, error_message is null. On failure, response is null and the error tells you why.

That last part matters in pipelines. A bad row does not have to crash the whole job. You can filter on error_message and handle the failures separately.

Turning answers into columns

Raw VARIANT is not what analysts want to see. They want columns.

Databricks lets you reach into a VARIANT with the colon path syntax. Here is a readable pattern for a support ticket table.

WITH decided AS (
  SELECT
    ticket_id,
    ticket_text,
    ai_decide(
      ticket_text,
      '{
        "team": {
          "type": "choice",
          "instructions": "Which team should handle this ticket?",
          "criteria": {
            "shipping": "Delivery or shipment issues",
            "billing": "Payment or invoice issues",
            "technical_support": "Product technical problems"
          }
        },
        "needs_escalation": {
          "type": "noul",
          "instructions": "Does this ticket need immediate escalation?"
        }
      }'
    ) AS decision
  FROM support_tickets
)

SELECT
  ticket_id,
  decision:response.answers.team.choice::STRING            AS team,
  decision:response.answers.team.confidence::DOUBLE        AS team_confidence,
  decision:response.answers.needs_escalation.probability::DOUBLE AS escalation_probability,
  decision:error_message::STRING                           AS error_message
FROM decided

Now you have a normal table. You can group by team, count escalations, join to customers, and build a dashboard. The AI part is just one column in a SQL query.

Where this fits in a real system

The most useful way to think about ai_decide is as a fast gate at the front of the pipeline.

            Incoming text
   (tickets, prompts, reviews, documents)
                   |
                   v
            [ ai_decide ]
       fast, structured, cheap
                   |
       +-----------+-----------+
       |                       |
  high confidence         low confidence
  clear decision          or high risk
       |                       |
       v                       v
  automatic route        larger model
  or normal pipeline     or a human

Most rows are easy. A clear billing ticket is a clear billing ticket. Those rows should be handled fast and cheaply.

A few rows are hard. Those deserve the expensive model, or a person.

ai_decide lets you tell the two apart, using the probabilities it already returns. That is the real design win. You stop paying premium prices for easy work.

Databricks names a few common uses for this pattern: routing prompts to the right model, tagging customer reviews with metadata, and evaluating agent quality, either in batch with SQL or in real time over REST (Databricks blog).

How it compares to the other AI Functions

Databricks already had AI Functions. So where does this one sit?

ai_query is the general tool. You write a prompt, pick a model, and get text or structured output back. It can do almost anything, including long reasoning and writing. That flexibility costs time and money.

ai_classify picks a label for text. It is a simpler, task-focused function.

ai_decide sits between them. It can ask several questions at once, mix yes or no, labels, and scales, and it returns probabilities and confidence for each. It does not write.

A simple rule of thumb:

Using confidence well

Probabilities are only useful if you act on them. Here is a calm, practical way to start.

Pick a threshold, then check it. For example, auto-route any ticket where team_confidence is above 0.85. Send the rest to a review queue.

Look at the review queue every week. If people keep agreeing with the model, raise the share it handles alone. If they keep correcting it, improve your criteria descriptions first.

Keep the decision with the data. Store the full VARIANT in a Delta table, not only the winning label. Months later, you will want to know how sure the model was when it made that call.

Expect small changes between runs. The documentation notes that generated answers can vary between calls. Do not build logic that needs the exact same probability every time. Build logic around ranges.

Governance is part of the point

The title of the Databricks announcement says "decisions on your governed data". That phrase is doing real work.

The documentation states that data is processed within the Databricks security perimeter, and that Databricks does not store the parameters passed into AI function calls (Databricks docs).

Because the function runs where your tables live, the normal rules still apply. Unity Catalog permissions decide who can read the source table. Lineage shows where the decision column came from. You are not exporting customer tickets to some outside service just to ask "billing or shipping?".

This is The Context Advantage idea in a small, practical form. Context comes from your own trusted tables. Control stays with your governance. Cost drops because the right tool does the right job. Choice stays open because the decision is a column, and the next step can be any model or any person you like.

Can I try this in Free Edition?

Be careful here.

ai_decide is a Beta feature that an admin must enable. It also depends on Databricks-managed model serving. Availability in Free Edition may be limited or may change. If the function is not available in your workspace, treat the SQL above as conceptual and use the exercise below to practise the pattern.

The pattern matters more than the function. Here it is in plain PySpark, runnable in Free Edition.

from pyspark.sql import functions as F

# Pretend these probabilities came from a decision model.
tickets_df = spark.createDataFrame(
    [
        (1, "My invoice is charged twice", "billing", 0.95),
        (2, "Package never arrived", "shipping", 0.91),
        (3, "App crashes and I think I was charged", "billing", 0.52),
        (4, "Login fails with error 500", "technical_support", 0.88),
    ],
    ["ticket_id", "ticket_text", "team", "team_confidence"],
)

CONFIDENCE_THRESHOLD = 0.85

routed_df = tickets_df.withColumn(
    "route",
    F.when(F.col("team_confidence") >= CONFIDENCE_THRESHOLD, F.col("team"))
     .otherwise(F.lit("human_review")),
)

display(routed_df)

Look at ticket 3. It mentions a crash and a charge. The confidence is low, so it goes to a person. That is exactly the kind of ticket you do not want auto-routed.

Try changing the threshold to 0.5 and run it again. Notice how quickly "automatic" turns into "risky". Choosing that line is a business decision, not a technical one.

What to remember

ai_decide is a small function with a big idea behind it. Not every AI task needs a model that writes.

Many of the questions inside a data pipeline are choices. Which bucket? How urgent? Yes or no? A decision model answers those faster, cheaper and in a cleaner shape than a generative model.

The skill for data engineers is not memorising the syntax. It is learning to see which questions in your system are really decisions, writing clear criteria for them, and deciding what happens when the model is not sure.

That is everyday data engineering work. It just has a new tool now.

Continue learning