Topic hub
Speech, vision, and multimodal AI
Models and tools that work with images, audio, speech, and video: image-text encoders, vision-language models, speech recognition, segmentation and vision backbones, and image and video generation, with the terms that come with them.
A multimodal model works with more than one kind of data, such as text, images, audio, or video. Some designs pair two encoders. OpenAI's CLIP trains an image encoder and a text encoder so that matching image and caption pairs score as similar, which lets it pick the most relevant text for an image without being trained for that specific task. Its model card describes CLIP as a research output: any deployed use, commercial or not, is currently out of scope; certain surveillance and facial recognition uses are always out of scope; and use should be limited to English.12
Vision-language models take images, and sometimes video, together with text and answer in text. Ai2's Molmo2-O-7B joins the SigLIP 2 vision encoder to Ai2's Olmo 3 7B Instruct language model and supports image, video, and multi-image understanding and grounding. Microsoft's Florence-2 uses simple text prompts to choose a task such as captioning, object detection, or segmentation. Google DeepMind's Gemma 4 12B Unified takes a different route: instead of separate encoders, it projects image patches and audio waveforms directly into the language model through lightweight linear layers, and it reads video as sequences of frames.345
Speech recognition models turn spoken audio into text. OpenAI's Whisper is a sequence-to-sequence model trained for multilingual speech recognition, translation of speech into English, and language identification. NVIDIA's parakeet-tdt-0.6b-v3 is a 600-million-parameter model that transcribes 25 European languages, detects the spoken language automatically, and adds punctuation and word-level timestamps. Whisper's model card warns that its output can include text that was not spoken, that accuracy varies across languages, accents, and dialects, and cautions against transcribing people recorded without their consent.678
Some vision models are building blocks for other software rather than chat systems. Meta's SAM 3 detects, segments, and tracks objects in images and video from short text phrases or from visual prompts such as points, boxes, and masks; SAM 3.1, released on March 27, 2026, added new checkpoints and a shared-memory approach to tracking several objects at once. Meta's DINOv3 is a family of vision foundation models that produce dense image features, and its repository includes code for semantic segmentation and depth estimation built on those features.911
Generative models create images, video, or audio. Hugging Face's Diffusers library for diffusion models is organized around three parts: pipelines that run a model in a few lines of code, interchangeable noise schedulers, and pretrained model components. Genmo's Mochi 1 preview is a 10-billion-parameter text-to-video diffusion model released under Apache 2.0. Its card says the initial release generates 480p video and that Genmo's reference code needs about 60 GB of GPU memory on a single GPU.1213
Terms are set for each release, so models of the same type can come with different conditions. parakeet-tdt-0.6b-v3 is under CC BY 4.0. Molmo2-O-7B is licensed under Apache 2.0, but its card says it is intended for research and educational use and was trained on third-party datasets subject to academic and non-commercial research use only. SAM 3 and DINOv3 weights are available only after an access request and come under Meta's own SAM License and DINOv3 License. Read the card and license for the specific checkpoint before you build on it.8391011
What this hub covers
Covers models and tools that take in or produce images, audio, speech, or video: image-text encoders, vision-language models, speech recognition, segmentation and vision backbones, and image and video generation. Robot control models are covered in the Robotics hub. It does not compare accuracy across models or report benchmark results. Featured records are examples chosen to cover the subject, not a ranking or a complete list.
Reading path
- Multimodal modelsStart here for how models handle images, audio, and video alongside text.
- How language models workVision-language models such as Molmo2 are built on a language model, so start there.
- How to read a model cardFind intended uses, out-of-scope uses, and known limitations for one release.
- What open weight and open source meanDownloadable weights are not the same as permission to use them for anything.
- Gated downloadSome vision checkpoints are released only after an access request.
- License scopeCode, weights, and training data can carry different terms.
- AI content provenanceHow the origin of generated images, audio, and video can be recorded and checked.
- Running AI on your own hardwareSome multimodal models, such as Gemma 4 12B, are documented for local use.
- RoboticsWhere vision and video models feed into robot control.
In the catalog
- Records that mention visionA text search that matches vision-language models, vision backbones, and segmentation models.
- Speech and audio recordsA text search that matches speech recognition models and the tools around them.
- Records that mention videoA text search that matches video understanding and video generation models.
- Multimodal model releasesModel records that mention multimodal input or output. Check each license before use.
Featured records
Examples chosen to cover the subject; not a ranking or a complete list.
- Ai2 (Allen Institute for AI)Organization
- MetaOrganization
- NVIDIAOrganization
- OpenAIOrganization
- GenmoOrganization
- CLIPOpen model or tool
- Molmo2-O 7BOpen model or tool
- Microsoft Florence-2Open model or tool
- Gemma 4 12B UnifiedOpen model or tool
- WhisperOpen model or tool
- NVIDIA parakeet-tdt-0.6b-v3Open model or tool
- SAM 3.1Open model or tool
- DINOv3Open model or tool
- DiffusersOpen model or tool
- Mochi 1 previewOpen model or tool
- EmbeddingGemma 2Open model or tool
Primary documents
- Model Card: CLIP (external site: raw.githubusercontent.com)OpenAI (GitHub)A short model card with direct statements of out-of-scope uses, a useful example of what to look for before deploying a vision model.
- Model Card: Whisper (external site: raw.githubusercontent.com)OpenAI (GitHub)Describes Whisper's training data by language and its documented limitations, including text that was not spoken and uneven accuracy across languages and accents.
- parakeet-tdt-0.6b-v3 model card (external site: huggingface.co)NVIDIA (Hugging Face)Lists supported languages, input format, audio length limits, and the license for one speech recognition release.
- Molmo2-O-7B model card (external site: huggingface.co)Ai2 (Hugging Face)Shows how a vision-language model names its vision encoder and language model, and how license and training-data terms can differ.
- Gemma 4 12B Unified (instruction-tuned) model card (external site: huggingface.co)Google (Hugging Face)Explains the encoder-free design of Gemma 4 12B and which input types each Gemma 4 size accepts.
- SAM 3: Segment Anything with Concepts (README) (external site: raw.githubusercontent.com)Meta (GitHub)SAM 3's description, the SAM 3.1 update, installation needs, and the access request required for checkpoints.
- DINOv3 README (external site: raw.githubusercontent.com)Meta (GitHub)Explains how to request DINOv3 weights and use the backbones and task adapters, with the DINOv3 License noted at the end.
- Diffusers README (external site: raw.githubusercontent.com)Hugging Face (GitHub)The library's own summary of its pipelines, schedulers, and model components.
- Mochi 1 preview model card (external site: huggingface.co)Genmo (Hugging Face)A video generation model card that states its hardware needs, output resolution, and known limitations.
Sources · reviewed Oct 8, 2026
Support Us
Help keep USASI useful.
Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).
About supporting this project