Annotate Video Support
Teaching it something new and it won't stick? Here are quick fixes and the most common questions.
Talk to the developer
Bug, idea or a thing it just won't learn. Every email gets read.
Troubleshooting
Train is greyed out
A model needs at least two labels with 5 clips each. The cat tells you how many clips are still missing. For better results, keep recording up to 25 or 60.
Training stops or fails
Keep the app open while it trains; it takes a few minutes and runs on your iPhone. If a label has clips that are too short to read frames from, record a few more. If it keeps failing, email us and mention which iPhone and iOS version you use.
Auto doesn't record anything
Check that a trained model is picked under Auto capture model, that Auto is on, and that camera access is allowed in Settings → Privacy & Security → Camera. Setting the certainty to 70 % makes it record more easily.
It records the wrong things
Answer ✗ on the “Was that …?” prompt to move a wrong clip to the “Not …” label, record a few more clips there, and train again. Setting the certainty to 95 % makes it stricter.
Where are my clips?
In the Files app under On My iPhone › Annotate Video, one folder per label.
Frequently asked questions
What does Annotate Video actually do?
It trains an image classifier, a small machine-learning model, from short video clips you record. Then it runs that model on your camera and records a clip whenever it recognises one of your labels. Recording, training and recognising all happen on your iPhone.
How many clips does it need?
Five clips each of two labels is enough to train a first model. Each clip gives the trainer several frames. The cat levels up at 5, 25 and 60 clips: 25 trains a solid model, 60 a sharp one.
Why do I need a “Not …” label?
A classifier always picks the closest label it knows. With a “Not Mug” label (your empty desk, other cups) it learns what isn't a mug, so it doesn't record every cup it sees. Auto only records the labels you want to spot, never the “Not …” ones.
Does it find where the thing is in the picture?
No. It's an image classifier, not an object detector: it looks at the whole camera image (a centred square) and says what it shows, with a confidence. That's why the thing you train should fill a good part of the frame.
How does it train on an iPhone?
With Apple's Create ML. It picks evenly spaced frames from every clip, keeps about one clip in five aside to test with, and trains in a few minutes. Afterwards it shows the accuracy on those unseen clips, the labels it mixes up, and clips it isn't sure about.
When does Auto record?
When the model is at least 70, 85 or 95 % sure, depending on what you set. Each clip starts 2 seconds before it spotted the label and keeps going while the label stays in view, up to 30 seconds.
Does it record sound?
No. Annotate Video only uses the camera, never the microphone. Clips are silent videos.
Does it watch in the background?
No, the camera only runs while the Capture tab is open. While Auto is on, the screen stays awake, so leave the iPhone on a stand or a charger.
Can I use my model and clips elsewhere?
Yes. Share a trained model as a Core ML file to use it in your own app, and find every clip in the Files app under On My iPhone › Annotate Video.
How is it related to Annotate Audio?
It's its twin. Annotate Audio trains a sound classifier and listens; Annotate Video trains an image classifier and watches. Same idea, same pixel look, a cat instead of a dog.
Requirements
iPhone with iOS 26.2 or later. Training runs on the iPhone itself.
See also the privacy policy.