Key takeaways
- A modality is a kind of data; a model that reads images or audio may still write only text, so check a card's input and output lists separately.
- Images and audio are usually converted into tokens, often by a separate encoder, before the language model reads them alongside text.
- Terms can differ by modality: Llama 4's use policy withholds license rights for its multimodal models from EU-based individuals and companies.
On this page
- Reading a card for inputs, outputs, and limits
- How pictures and sound become tokens
- Speech recognition and text-to-speech
- Image and video generation by diffusion
- Licenses and use policies can differ by modality
- Privacy with photos and voices
- Worked example: one project, three modalities
- What you can do next
- Check your understanding
A modality is a kind of data, such as text, images, audio, or video. A text-only model reads and writes text; a multimodal model also accepts or produces at least one other kind. Models that "see" or "hear" usually convert a picture or sound clip into the kind of internal pieces they use for text, often with a separate part called an encoder, then answer in text. Reading a modality is not producing it: a model that describes photos may be unable to create one. Licenses, use policies, and privacy risks can also differ by modality. A release's model card, the publisher's documentation for it, says which inputs and outputs it supports and where its limits are.
Reading a card for inputs, outputs, and limits
Model cards usually list what a model accepts and what it returns as separate items:
- Meta's Llama 4 model card (external site: github.com) lists input as multilingual text and image, and output as multilingual text and code.
- Google's Gemma 4 model card (external site: ai.google.dev) says every size accepts text and images, the E2B, E4B, and 12B sizes also accept audio, and all of them generate text.
- OpenAI's Whisper model card (external site: github.com) describes audio in and text out: transcription, plus translation into English.
- NVIDIA's MagpieTTS card (external site: huggingface.co) is the reverse: text in, audio out.
- Microsoft's Florence-2-large card (external site: huggingface.co) takes an image plus a short task prompt, such as
<CAPTION>or<OD>for object detection, and returns text, including box coordinates for detected objects.
Cards also state limits for each modality. Gemma 4 accepts up to 30 seconds of audio and 60 seconds of video at one frame per second; Llama 4 was tested with up to five input images; and Parakeet expects 16 kHz mono audio.
How pictures and sound become tokens
Language models work on tokens, small pieces of text represented as numbers (see tokens and context windows). Other kinds of input have to be converted into something similar.
- Images. Ai2's Molmo2 paper (external site: arxiv.org) describes a common design. Images are split or resized into fixed-size crops; a vision encoder, a network trained on images, turns each crop into features for small patches; and a connector pools those features and projects them into "visual tokens" that go to the language model along with the text. The Molmo2-O-7B card (external site: huggingface.co) pairs Ai2's Olmo 3 7B Instruct language model with a SigLIP vision backbone published by Google. In Gemma 4 31B the vision encoder has about 550 million parameters, and developers can choose how many tokens an image uses, with budgets of 70, 140, 280, 560, or 1,120; more tokens keep more detail but take more computation.
- Other designs. Meta's Llama 4 announcement (external site: ai.meta.com) describes "early fusion," which feeds text and vision tokens into one model backbone, with a vision encoder based on MetaCLIP. Gemma 4 12B Unified has no separate encoders at all: its card says it projects raw image patches and audio waveforms directly into the language model's embedding space through lightweight linear layers.
- Audio. The Whisper paper (external site: arxiv.org) resamples audio to 16,000 samples per second and converts it into a log-Mel spectrogram, a grid showing how strong each frequency band is in each short slice of time, computed on 25-millisecond windows every 10 milliseconds. An encoder-decoder Transformer reads that grid and writes text.
Speech recognition and text-to-speech
Automatic speech recognition (ASR) turns speech into text. The Whisper model card says the models were trained on 680,000 hours of internet audio with transcripts, perform unevenly across languages and accents, and may output text that was never spoken (hallucination). NVIDIA's parakeet-tdt-0.6b-v3 (external site: huggingface.co) is a 600-million-parameter ASR model for 25 European languages that returns text with punctuation, capitalization, and timestamps.
Text-to-speech (TTS) runs the other way. MagpieTTS predicts discrete audio codec tokens, compressed units of sound that a decoder turns back into a waveform, and outputs 22.05 kHz WAV audio in five built-in voices across 12 languages. Its card says it is not intended for zero-shot voice cloning (copying a new voice from a short sample). OpenAI's text-to-speech guide (external site: developers.openai.com) says its usage policies require telling end users that the voice they hear is AI-generated.
Image and video generation by diffusion
Some image and video generators use diffusion. The 2020 Denoising Diffusion Probabilistic Models paper (external site: arxiv.org) by UC Berkeley researchers describes a diffusion process that gradually adds noise to data until the signal is destroyed; the model learns to reverse it, so generating starts from noise and removes it step by step. Genmo's Mochi 1 preview card (external site: huggingface.co), tagged text-to-video, describes a 10-billion-parameter diffusion model that encodes the text prompt with a T5-XXL language model and works on video compressed by a separate autoencoder, a network that shrinks data and rebuilds it. Its card lists limits: 480p output, some warping with extreme motion, and poor results on animated content.
Licenses and use policies can differ by modality
- The Llama 4 Acceptable Use Policy (external site: dev.meta.ai) says that for any multimodal models in Llama 4, the license rights in Section 1(a) are not granted to individuals domiciled in, or companies with a principal place of business in, the European Union. The restriction does not apply to end users of a product that incorporates those models. The Llama 4 card lists image input for both Scout and Maverick.
- One publisher may use different terms for models with different inputs and outputs. NVIDIA's Parakeet card uses CC BY 4.0 (external site: creativecommons.org), which requires attribution when you share the licensed material; the MagpieTTS card uses the NVIDIA Open Model License.
- Hosted services add their own rules. Besides the TTS disclosure requirement, OpenAI's custom voices guide (external site: developers.openai.com) limits the feature to eligible customers and requires a recorded consent statement from the voice actor.
Privacy with photos and voices
- Location. Apple's guide to location metadata (external site: support.apple.com) explains that when Location Services is on for the Camera app, the coordinates where a photo or video was taken are embedded in it, and people you share it with may be able to see them. The guide shows how to remove them.
- Other people. The Whisper card cautions against transcribing recordings of people made without their consent, and says the models may be able to recognize specific individuals.
- Inferences. The Llama 4 Acceptable Use Policy prohibits collecting or inferring private or sensitive information about individuals, such as identity, health, or demographic information, unless you have the right to do so under applicable law.
- Where it runs. A hosted service receives your files; local inference keeps them on your hardware (hosted or local, AI and your data).
Worked example: one project, three modalities
Morgan is a fictional reader invented for this page. Morgan volunteers for a local history society that has recorded interviews, made with signed consent, and scanned photos with handwritten captions. The society wants transcripts, photo descriptions for its website, and short narrated summaries.
- Split the job by modality. Speech to text, image to text, and text to speech are three jobs, so Morgan reads each candidate card's input and output lines instead of assuming one model does everything.
- Transcripts. Parakeet v3 supports English and expects 16 kHz mono audio, so Morgan converts the recordings first. Because the Whisper card warns of text that was never spoken, a volunteer checks every transcript against its recording.
- Photo descriptions. Gemma 4 31B accepts images and returns text. Its card lists handwriting recognition and suggests higher image token budgets for reading small text, so Morgan uses one.
- Narration. MagpieTTS offers built-in voices and is not meant for voice cloning, so Morgan does not try to copy an interviewee's voice, and the society labels the audio as AI-generated.
- Terms. Morgan records each license: CC BY 4.0 for Parakeet, Apache 2.0 for Gemma 4, and the NVIDIA Open Model License for MagpieTTS.
Morgan's outcome: three tools, each checked against its own card, run locally because the recordings name living people. A hosted service could work once the society approves its data policy.
What you can do next
- Open catalog records for Parakeet, Whisper, Florence-2, Molmo, Gemma 4 31B, Llama 4 Scout, and Mochi 1 preview.
- Read how to read a model card and how language models work.
- Look up model card, tokenizer, and acceptable use policy in the glossary.
- Browse the foundation models and local AI hubs.
Sources
All read on October 8, 2026.
- Meta: Llama 4 model card (external site: github.com), Llama 4 Acceptable Use Policy (external site: dev.meta.ai), The Llama 4 herd (external site: ai.meta.com) (April 5, 2025)
- Google: Gemma 4 model card (external site: ai.google.dev)
- OpenAI: Whisper model card (external site: github.com), Robust Speech Recognition via Large-Scale Weak Supervision (external site: arxiv.org) (arXiv 2212.04356), Text to speech guide (external site: developers.openai.com), Custom voices guide (external site: developers.openai.com)
- NVIDIA: parakeet-tdt-0.6b-v3 model card (external site: huggingface.co), magpie_tts_multilingual_357m model card (external site: huggingface.co)
- Ai2: Molmo2 paper (external site: arxiv.org) (arXiv 2601.10611), Molmo2-O-7B model card (external site: huggingface.co)
- Microsoft: Florence-2-large model card (external site: huggingface.co)
- Genmo: mochi-1-preview model card (external site: huggingface.co)
- UC Berkeley researchers: Denoising Diffusion Probabilistic Models (external site: arxiv.org) (arXiv 2006.11239)
- Creative Commons: Attribution 4.0 International legal code (external site: creativecommons.org)
- Apple: Manage location metadata in Photos (external site: support.apple.com)
Check your understanding
Three quick questions, answered from this page. Nothing you choose is saved or sent anywhere.
0 of 3 answered.