d1-3B
Liquid AI's zero-shot classification model for text and images, with a guide to running it on Jetson using PyTorch.
Model Details
Get started with d1-3B on Jetson
d1-3B is a 3B zero-shot classification model from Liquid AI, built on LFM2.5-VL-3B. Given text, images, and a set of questions, it returns structured answers in a single forward pass.
Describe your task and candidate labels in natural language to route support tickets, classify intent, moderate content, or inspect images. Changing tasks requires changing the questions and labels, without task-specific fine-tuning.
The model answers three types of questions: noul returns a probability for a yes/no question, choice selects from named options, and score returns a rating across ordered levels. Several questions can share the same input. The model returns decisions directly without generating text.
This model answers a single question in 16 ms on Jetson AGX Thor, 26 ms on Jetson AGX Orin 64 GB, and 50 ms on Jetson Orin Nano after warmup in Liquid’s benchmarks. On AGX Thor, three questions about the same input take 20 ms in one pass.
See Liquid’s Open d1 Arcade for demos of camera games, drawing recognition, and message classification to inspire your own applications.
This guide uses the BF16 checkpoint to run inference on Jetson. For a smaller model that also accepts audio, see d1-omni-600M on Jetson.
1. Prerequisites
| Requirement | Configuration |
|---|---|
| Hardware | Jetson AGX Thor, Jetson AGX Orin 64 GB, or Jetson Orin Nano |
| System software | JetPack 7.2 / Jetson Linux 39.2 |
| Container | Docker installed and NVIDIA Container Toolkit configured for GPU access |
2. Start the container
NVIDIA’s PyTorch container includes PyTorch built for the Jetson GPU.
docker run --pull always --rm -it --runtime nvidia --ipc=host \
-v "$PWD":/workspace -w /workspace \
-v ~/.cache/huggingface:/root/.cache/huggingface \
nvcr.io/nvidia/pytorch:26.09-py3
Run the remaining commands inside the container. Files in /workspace and downloaded models stay on the host; installed packages disappear when this temporary container exits.
3. Install the dependencies
pip install "transformers==5.18.0" pillow
The remaining dependencies are already installed and configured in the container.
4. Run inference and check GPU execution
from_pretrained() downloads the checkpoint on the first run and reuses the mounted Hugging Face cache on later runs.
Save this as example.py. It is adapted from the “How to use” example in the d1-3B model card.
import torch
from transformers import AutoModel
from transformers.image_utils import load_image
assert torch.cuda.is_available(), "CUDA is not available"
model = AutoModel.from_pretrained("LiquidAI/d1-3B", trust_remote_code=True, dtype=torch.bfloat16).to("cuda")
# Text: several named questions over one state, answered in one pass
questions = {
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?",
},
"team": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Charges, refunds, invoices",
"technical": "App or site faults",
"fraud": "Suspected unauthorised use",
},
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": ["Can wait", "Today", "Blocking the customer now"],
},
}
text_result = model.system_one("I was charged twice this month, please refund one of them.", questions)
# Image: the photo is the whole state
image = load_image("http://images.cocodataset.org/val2017/000000039769.jpg") # two cats on a sofa
cats = {
"type": "choice",
"instructions": "How many cats are there?",
"criteria": {"one": "One", "two": "Two", "more": "Three or more"},
}
image_result = model.system_one(None, {"cats": cats}, images=[image])
# Batch: many requests, packed together with no padding
tickets = ["Where is my parcel? It was due Monday.", "The app crashes when I open settings."]
batch_result = model.system_one_batch([(t, {"team": questions["team"]}) for t in tickets])
print("GPU:", torch.cuda.get_device_name())
print("Refund:", f'{text_result["answers"]["refund"]["noul"]:.2f}')
print("Team:", text_result["answers"]["team"]["choice"])
print("Urgency:", f'{text_result["answers"]["urgency"]["score"]:.2f}')
print("Cats:", image_result["answers"]["cats"]["choice"])
print("Batch:", ", ".join(r["answers"]["team"]["choice"] for r in batch_result))
python example.py
Example output from a Jetson Thor run of the model-card example, formatted to match the print statements above:
GPU: NVIDIA Thor
Refund: 0.99
Team: billing
Urgency: 0.76
Cats: two
Batch: technical, technical
Values can differ between devices. The GPU name should match your Jetson.