Meta's efficient 1B/3B and vision-capable 11B/90B open models for edge and multimodal use.
Llama 3.2 extended the family with tiny 1B and 3B text models that run on phones and laptops, plus 11B and 90B vision-language models that can read images. It's the practical choice when you need on-device inference or cheap multimodal understanding without a cloud bill.
Who it's for: Mobile and edge developers, and teams needing cheap vision understanding on private infrastructure.
Runs offline on phones and consumer hardware.
11B/90B variants understand images and charts.
Self-host or fine-tune freely.
Tiny sizes mean instant local responses.
Llama 3.2 is the smart pick for edge and private multimodal apps. Use the 90B vision model when you need image understanding without a vendor API.