Independent project. Not a U.S. government website.

USASI

Topic hub

Speech, vision, and multimodal AI

Models and tools that work with images, audio, speech, and video: image-text encoders, vision-language models, speech recognition, segmentation and vision backbones, and image and video generation, with the terms that come with them.

A multimodal model works with more than one kind of data, such as text, images, audio, or video. Some designs pair two encoders. OpenAI's CLIP trains an image encoder and a text encoder so that matching image and caption pairs score as similar, which lets it pick the most relevant text for an image without being trained for that specific task. Its model card describes CLIP as a research output: any deployed use, commercial or not, is currently out of scope; certain surveillance and facial recognition uses are always out of scope; and use should be limited to English.12

Vision-language models take images, and sometimes video, together with text and answer in text. Ai2's Molmo2-O-7B joins the SigLIP 2 vision encoder to Ai2's Olmo 3 7B Instruct language model and supports image, video, and multi-image understanding and grounding. Microsoft's Florence-2 uses simple text prompts to choose a task such as captioning, object detection, or segmentation. Google DeepMind's Gemma 4 12B Unified takes a different route: instead of separate encoders, it projects image patches and audio waveforms directly into the language model through lightweight linear layers, and it reads video as sequences of frames.345

Speech recognition models turn spoken audio into text. OpenAI's Whisper is a sequence-to-sequence model trained for multilingual speech recognition, translation of speech into English, and language identification. NVIDIA's parakeet-tdt-0.6b-v3 is a 600-million-parameter model that transcribes 25 European languages, detects the spoken language automatically, and adds punctuation and word-level timestamps. Whisper's model card warns that its output can include text that was not spoken, that accuracy varies across languages, accents, and dialects, and cautions against transcribing people recorded without their consent.678

Some vision models are building blocks for other software rather than chat systems. Meta's SAM 3 detects, segments, and tracks objects in images and video from short text phrases or from visual prompts such as points, boxes, and masks; SAM 3.1, released on March 27, 2026, added new checkpoints and a shared-memory approach to tracking several objects at once. Meta's DINOv3 is a family of vision foundation models that produce dense image features, and its repository includes code for semantic segmentation and depth estimation built on those features.911

Generative models create images, video, or audio. Hugging Face's Diffusers library for diffusion models is organized around three parts: pipelines that run a model in a few lines of code, interchangeable noise schedulers, and pretrained model components. Genmo's Mochi 1 preview is a 10-billion-parameter text-to-video diffusion model released under Apache 2.0. Its card says the initial release generates 480p video and that Genmo's reference code needs about 60 GB of GPU memory on a single GPU.1213

Terms are set for each release, so models of the same type can come with different conditions. parakeet-tdt-0.6b-v3 is under CC BY 4.0. Molmo2-O-7B is licensed under Apache 2.0, but its card says it is intended for research and educational use and was trained on third-party datasets subject to academic and non-commercial research use only. SAM 3 and DINOv3 weights are available only after an access request and come under Meta's own SAM License and DINOv3 License. Read the card and license for the specific checkpoint before you build on it.8391011

What this hub covers

Covers models and tools that take in or produce images, audio, speech, or video: image-text encoders, vision-language models, speech recognition, segmentation and vision backbones, and image and video generation. Robot control models are covered in the Robotics hub. It does not compare accuracy across models or report benchmark results. Featured records are examples chosen to cover the subject, not a ranking or a complete list.

Reading path

  1. 1.Multimodal modelsStart here for how models handle images, audio, and video alongside text.
  2. 2.How language models workVision-language models such as Molmo2 are built on a language model, so start there.
  3. 3.How to read a model cardFind intended uses, out-of-scope uses, and known limitations for one release.
  4. 4.What open weight and open source meanDownloadable weights are not the same as permission to use them for anything.
  5. 5.Gated downloadSome vision checkpoints are released only after an access request.
  6. 6.License scopeCode, weights, and training data can carry different terms.
  7. 7.AI content provenanceHow the origin of generated images, audio, and video can be recorded and checked.
  8. 8.Running AI on your own hardwareSome multimodal models, such as Gemma 4 12B, are documented for local use.
  9. 9.RoboticsWhere vision and video models feed into robot control.

In the catalog

Examples chosen to cover the subject; not a ranking or a complete list.

Primary documents

Sources · reviewed Oct 8, 2026

  1. 1.
    CLIP README (external site: raw.githubusercontent.com)

    OpenAI (GitHub) · Repository · accessed Oct 8, 2026

  2. 2.
    Model Card: CLIP (external site: raw.githubusercontent.com)

    OpenAI (GitHub) · Model card · accessed Oct 8, 2026

  3. 3.
    Molmo2-O-7B model card (external site: huggingface.co)

    Ai2 (Hugging Face) · Model card · accessed Oct 8, 2026

  4. 4.
    Florence-2-large model card (external site: huggingface.co)

    Microsoft (Hugging Face) · Model card · accessed Oct 8, 2026

  5. 5.
    Gemma 4 12B Unified (instruction-tuned) model card (external site: huggingface.co)

    Google (Hugging Face) · Model card · accessed Oct 8, 2026

  6. 6.
    Whisper README (external site: raw.githubusercontent.com)

    OpenAI (GitHub) · Repository · accessed Oct 8, 2026

  7. 7.
    Model Card: Whisper (external site: raw.githubusercontent.com)

    OpenAI (GitHub) · Model card · accessed Oct 8, 2026

  8. 8.
    parakeet-tdt-0.6b-v3 model card (external site: huggingface.co)

    NVIDIA (Hugging Face) · Model card · accessed Oct 8, 2026

  9. 9.
  10. 10.
    SAM License (external site: raw.githubusercontent.com)

    Meta (GitHub) · License · published Nov 19, 2025 · accessed Oct 8, 2026

  11. 11.
    DINOv3 README (external site: raw.githubusercontent.com)

    Meta (GitHub) · Repository · accessed Oct 8, 2026

  12. 12.
    Diffusers README (external site: raw.githubusercontent.com)

    Hugging Face (GitHub) · Repository · accessed Oct 8, 2026

  13. 13.
    Mochi 1 preview model card (external site: huggingface.co)

    Genmo (Hugging Face) · Model card · accessed Oct 8, 2026

Support Us

Help keep USASI useful.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project