Segment anything from text
Describe any object in natural language and get a precise mask.
Segment anything with text and visual prompts.
Use language and visual examples to precisely identify, segment, and follow any object in images or videos.

INTERACTIVE DEMO
Try a text prompt or add a visual example to explore SAM 3.
SAM 3 IN ACTION
TEXT: “PERSON” · 98%
Find concepts beyond a fixed label set.

Separate every matching object at once.

Follow concepts as scenes change.

Explore fine-grained real-world objects.
KEY CAPABILITIES
Describe any object in natural language and get a precise mask.
Point to an example image when language alone is not enough.
Find all matching objects in a crowded scene.
Track concepts through motion, occlusion, and camera changes.
Handle long-tail, detailed, and real-world categories.
Turn open-vocabulary prompts into scalable annotations.
HOW IT WORKS
Describe a concept or show an example.
Get a precise mask for every instance.
Follow the concept through your video.
CHOOSE YOUR WORKFLOW
USE CASES
SAM 3 WORKSPACE
Start with a flexible pack and use every SAM 3 workflow.
$9 one time
$19 one time
$39 one time
$79 one time
FAQ
SAM 3 is a unified model that uses text and visual prompts to identify, segment, and follow objects in images and videos.
Use natural language, a visual exemplar, or both together to define the concept you want to find.
Yes. SAM 3 can detect and segment every matching instance in an image or video.
Yes. It follows prompted concepts across frames while maintaining consistent masks.
Creators, researchers, robotics teams, dataset builders, and anyone working with visual data.
START EXPLORING