Qwen3.5 4B
Alibaba's efficient Qwen3.5 4B vision-language model tuned for practical multimodal deployment
Serve the model
Start server
Choose module, then engine and optional parameters on the left, then copy the serve command by clicking the button on the right.
Command
ยท
No command for this module and engine in model data.
Call the model over Web API
Copy a client command below and paste it into your terminal to make a Web API request to the model you just served.
curl -s http://${JETSON_HOST}:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "cyankiwi/Qwen3.5-4B-AWQ-4bit",
"messages": [{"role": "user", "content": "Hello!"}]
}' Benchmark
Qwen3.5-4B · vLLM · NVFP4 / W4A16 · ISL 2048 / OSL 128
C = concurrent requests. Results will vary with image, clocks, and workload.
Model Details
Qwen3.5 4B offers a balanced point in the Qwen3.5 family for local multimodal instruction following, visual understanding, and agent-style workloads on Jetson.
Inputs and Outputs
Input: Text and images
Output: Text
Intended Use Cases
- Visual question answering: Multimodal prompting with image inputs
- Image understanding: Captioning, scene analysis, and grounded responses
- Tool calling: Structured tool use with vLLM
- Multilingual tasks: Translation and multilingual prompting
Additional Resources
- Original Model - Base Qwen3.5 4B checkpoint
- AWQ Checkpoint - Quantized checkpoint used here