Independent project. Not a U.S. government website.

USASI

Explainer · AI basics

Multimodal models: text, images, audio, and video

How do models that see and hear differ from text-only models?

Intermediate6 min readReviewed Oct 8, 2026

General information, not legal or professional advice. All explainers

How this page was made
  • Researched and written with AI assistance from primary sources, which are listed at the end with the date they were read.
  • Source-checked: a separate AI fact-check pass compared each sentence with its source and corrected what did not match (fact-check report (external site: github.com)).
  • Automated checks passed: links, structure, and formatting are validated before every publish.
  • Not individually reviewed by a person before publication. What these review levels mean

Key takeaways

  • A modality is a kind of data; a model that reads images or audio may still write only text, so check a card's input and output lists separately.
  • Images and audio are usually converted into tokens, often by a separate encoder, before the language model reads them alongside text.
  • Terms can differ by modality: Llama 4's use policy withholds license rights for its multimodal models from EU-based individuals and companies.
On this page

A modality is a kind of data, such as text, images, audio, or video. A text-only model reads and writes text; a multimodal model also accepts or produces at least one other kind. Models that "see" or "hear" usually convert a picture or sound clip into the kind of internal pieces they use for text, often with a separate part called an encoder, then answer in text. Reading a modality is not producing it: a model that describes photos may be unable to create one. Licenses, use policies, and privacy risks can also differ by modality. A release's model card, the publisher's documentation for it, says which inputs and outputs it supports and where its limits are.

Reading a card for inputs, outputs, and limits

Model cards usually list what a model accepts and what it returns as separate items:

Cards also state limits for each modality. Gemma 4 accepts up to 30 seconds of audio and 60 seconds of video at one frame per second; Llama 4 was tested with up to five input images; and Parakeet expects 16 kHz mono audio.

How pictures and sound become tokens

Language models work on tokens, small pieces of text represented as numbers (see tokens and context windows). Other kinds of input have to be converted into something similar.

  • Images. Ai2's Molmo2 paper (external site: arxiv.org) describes a common design. Images are split or resized into fixed-size crops; a vision encoder, a network trained on images, turns each crop into features for small patches; and a connector pools those features and projects them into "visual tokens" that go to the language model along with the text. The Molmo2-O-7B card (external site: huggingface.co) pairs Ai2's Olmo 3 7B Instruct language model with a SigLIP vision backbone published by Google. In Gemma 4 31B the vision encoder has about 550 million parameters, and developers can choose how many tokens an image uses, with budgets of 70, 140, 280, 560, or 1,120; more tokens keep more detail but take more computation.
  • Other designs. Meta's Llama 4 announcement (external site: ai.meta.com) describes "early fusion," which feeds text and vision tokens into one model backbone, with a vision encoder based on MetaCLIP. Gemma 4 12B Unified has no separate encoders at all: its card says it projects raw image patches and audio waveforms directly into the language model's embedding space through lightweight linear layers.
  • Audio. The Whisper paper (external site: arxiv.org) resamples audio to 16,000 samples per second and converts it into a log-Mel spectrogram, a grid showing how strong each frequency band is in each short slice of time, computed on 25-millisecond windows every 10 milliseconds. An encoder-decoder Transformer reads that grid and writes text.

Speech recognition and text-to-speech

Automatic speech recognition (ASR) turns speech into text. The Whisper model card says the models were trained on 680,000 hours of internet audio with transcripts, perform unevenly across languages and accents, and may output text that was never spoken (hallucination). NVIDIA's parakeet-tdt-0.6b-v3 (external site: huggingface.co) is a 600-million-parameter ASR model for 25 European languages that returns text with punctuation, capitalization, and timestamps.

Text-to-speech (TTS) runs the other way. MagpieTTS predicts discrete audio codec tokens, compressed units of sound that a decoder turns back into a waveform, and outputs 22.05 kHz WAV audio in five built-in voices across 12 languages. Its card says it is not intended for zero-shot voice cloning (copying a new voice from a short sample). OpenAI's text-to-speech guide (external site: developers.openai.com) says its usage policies require telling end users that the voice they hear is AI-generated.

Image and video generation by diffusion

Some image and video generators use diffusion. The 2020 Denoising Diffusion Probabilistic Models paper (external site: arxiv.org) by UC Berkeley researchers describes a diffusion process that gradually adds noise to data until the signal is destroyed; the model learns to reverse it, so generating starts from noise and removes it step by step. Genmo's Mochi 1 preview card (external site: huggingface.co), tagged text-to-video, describes a 10-billion-parameter diffusion model that encodes the text prompt with a T5-XXL language model and works on video compressed by a separate autoencoder, a network that shrinks data and rebuilds it. Its card lists limits: 480p output, some warping with extreme motion, and poor results on animated content.

Licenses and use policies can differ by modality

  • The Llama 4 Acceptable Use Policy (external site: dev.meta.ai) says that for any multimodal models in Llama 4, the license rights in Section 1(a) are not granted to individuals domiciled in, or companies with a principal place of business in, the European Union. The restriction does not apply to end users of a product that incorporates those models. The Llama 4 card lists image input for both Scout and Maverick.
  • One publisher may use different terms for models with different inputs and outputs. NVIDIA's Parakeet card uses CC BY 4.0 (external site: creativecommons.org), which requires attribution when you share the licensed material; the MagpieTTS card uses the NVIDIA Open Model License.
  • Hosted services add their own rules. Besides the TTS disclosure requirement, OpenAI's custom voices guide (external site: developers.openai.com) limits the feature to eligible customers and requires a recorded consent statement from the voice actor.

Privacy with photos and voices

  • Location. Apple's guide to location metadata (external site: support.apple.com) explains that when Location Services is on for the Camera app, the coordinates where a photo or video was taken are embedded in it, and people you share it with may be able to see them. The guide shows how to remove them.
  • Other people. The Whisper card cautions against transcribing recordings of people made without their consent, and says the models may be able to recognize specific individuals.
  • Inferences. The Llama 4 Acceptable Use Policy prohibits collecting or inferring private or sensitive information about individuals, such as identity, health, or demographic information, unless you have the right to do so under applicable law.
  • Where it runs. A hosted service receives your files; local inference keeps them on your hardware (hosted or local, AI and your data).

Worked example: one project, three modalities

Morgan is a fictional reader invented for this page. Morgan volunteers for a local history society that has recorded interviews, made with signed consent, and scanned photos with handwritten captions. The society wants transcripts, photo descriptions for its website, and short narrated summaries.

  1. Split the job by modality. Speech to text, image to text, and text to speech are three jobs, so Morgan reads each candidate card's input and output lines instead of assuming one model does everything.
  2. Transcripts. Parakeet v3 supports English and expects 16 kHz mono audio, so Morgan converts the recordings first. Because the Whisper card warns of text that was never spoken, a volunteer checks every transcript against its recording.
  3. Photo descriptions. Gemma 4 31B accepts images and returns text. Its card lists handwriting recognition and suggests higher image token budgets for reading small text, so Morgan uses one.
  4. Narration. MagpieTTS offers built-in voices and is not meant for voice cloning, so Morgan does not try to copy an interviewee's voice, and the society labels the audio as AI-generated.
  5. Terms. Morgan records each license: CC BY 4.0 for Parakeet, Apache 2.0 for Gemma 4, and the NVIDIA Open Model License for MagpieTTS.

Morgan's outcome: three tools, each checked against its own card, run locally because the recordings name living people. A hosted service could work once the society approves its data policy.

What you can do next

Sources

All read on October 8, 2026.

Check your understanding

Three quick questions, answered from this page. Nothing you choose is saved or sent anywhere.

  1. 1.A model card says a model accepts photos and returns text descriptions. What can you conclude about whether it can create images?
  2. 2.In the common image design the page describes from the Molmo2 paper, how does a picture reach the language model?
  3. 3.NVIDIA publishes both Parakeet (speech to text) and MagpieTTS (text to speech). What does the page say about their licenses?

0 of 3 answered.

Keep learning

Part of New to AI.

Support Us

Help keep USASI useful.

Find the catalog useful? Leave an optional tip to support its upkeep. Tips never affect listings, coverage, or openness assessments.

Optional. No USASI account required. Payment takes place on the linked provider’s website (Buy Me a Coffee).

About supporting this project