หน้านี้ตอบอะไร
Llama 3.2 11B Vision is a multimodal model with 11 billion parameters, designed to handle tasks combining visual and textual data. It excels in tasks such as image captioning and...
- Meta · meta-llama/llama-3.2-11b-vision-instruct
- text+image->text · เส้นทางโมเดลทั่วโลก
- 131,072 context · US$0.245 input