What a Video Call Feature Really Involves
The word call implies two cameras and two people. In this category it usually means one: something is rendered on your screen while you talk, and often nothing is captured from your side at all. What is rendered can be a looping clip, a rigged avatar being lip-synced, generated frames, or a still image with an animation on top — four quite different products behind one label.
So there are two questions, and the marketing page answers neither. What exactly is on the screen, and is your camera involved.
The four things that might be on screen
A pre-rendered loop. A short clip of the character idling, played on repeat while the conversation happens elsewhere. Cheap, reliable, and recognisable by repetition — watch for a gesture or blink pattern that recurs on a fixed cycle.
A rigged avatar, lip-synced. A 2D or 3D model whose mouth is driven by the audio and whose expressions come from a small library of animations. This is the most common genuine implementation. It reacts to the shape of the speech rather than to the meaning, which is why the expression sometimes does not match what is being said.
Generated frames. A talking face produced rather than animated. More convincing when it works, expensive to run, and the most likely to be limited in length or availability.
A still image with motion effects. A photograph or illustration with a subtle drift, a waveform, or a glow that responds to the audio. Legitimate as a design choice, and it is not video in any sense a reader would expect from the word.
Telling them apart takes a minute of deliberate watching. Say something highly specific and see whether anything on screen responds to the content rather than to the sound. Sit in silence and watch the idle behaviour for a loop. Look at whether the eyes ever go anywhere unexpected.
Whether your camera is on
This is the part worth being deliberate about, because a two-way version is a materially different product with a materially different footprint.
Most implementations are one-way. You hear and see; nothing of you is captured. If the app never requests camera access, that is settled.
Some are genuinely two-way, and use the front camera to drive a reaction — a nod when you nod, a comment on what you are holding up. That requires the camera, and it means frames are being handled somewhere.
The question that follows is where they go. A reaction driven on the device is a different thing from frames uploaded for processing, and the observable difference is data usage and whether the feature works with the network off. What airplane mode tells you about an app is the relevant test, and what on-device processing would look like covers what to expect from the claim.
Granting the camera is not the same as knowing what is done with it. The scope of a prompt and the scope of the use are separate, as set out in what a permission prompt actually grants. If a video feature works without camera access, leaving it off is the simple answer.
What video costs that voice does not
It requires your eyes. Voice can be used while walking, cooking, or lying in the dark. Video cannot, which narrows the situations the feature is any use in more than people expect when they choose an app for it.
It is visible over your shoulder. A screen showing a face is legible from further away than a screen showing text, and less deniable at a glance.
It moves real data and real battery. Both matter on a metered connection or a long day.
It is the most expensive feature to operate, which is why it is the one most often restricted, queued, shortened, or withdrawn. A feature that is costly to run is a feature whose terms change, and that pattern is covered in what an app update can change without asking.
Recording it makes a file. A screenshot or screen recording of a call becomes an item in your photo library with everything that follows from that, described in what a screenshot becomes after you take it.
Why quality claims age badly here
Rendered faces sit in the part of the quality curve where improvements are dramatic and constant. Any specific judgement about how convincing this looks is out of date quickly, which is a reason to evaluate the structure of the feature rather than its current fidelity.
The structural questions do not age. Is it a loop or is it responsive. Is my camera involved. Does it work offline. Is it length-limited. Is it available on the terms I signed up under. Those answers stay useful after the visuals have moved on.
The limits of watching
You cannot tell from the screen whether frames from your camera were uploaded, retained, or discarded, and no amount of careful observation will establish it. Data usage and offline behaviour narrow the possibilities; they do not close the question, and treating them as proof would be overreach.
What you can settle is the shape of what you are buying: whether the call is a conversation with something responsive or a video playing behind a chat, and whether you are on camera. Both are answerable in a few minutes, and both are more useful than any impression of how lifelike it looked.