1 link tagged with all of: deep-learning + multimodal + ai + model + vision-language
Links
Youtu-VL is a 4B-parameter Vision-Language Model that excels in both vision-centric and general multimodal tasks without needing task-specific modules. It uses a unique autoregressive supervision method to enhance visual understanding and preserve detailed information. The model supports various applications, from image classification to visual question answering.
vision-language ✓
multimodal ✓
model ✓
deep-learning ✓
ai ✓