Does a Visual Companion Need a Different Evaluation?

Apps described as virtual friends or virtual companions usually differ from text-only ones in presentation: there is a face, an avatar, sometimes a scene, sometimes lip-synced speech. That difference adds exactly three things to check and changes nothing about how you evaluate the conversation, because the conversation is produced the same way either way. This post is deliberately short, because that is the whole finding.

The general evaluation sequence — what to check before installing, in the first session, before paying, and after a few weeks — is set out in how to evaluate an AI friend app before you pay for one. Everything there applies unchanged. What follows is the supplement.

The three additions

Each of these exists because a visual layer needs things a text field does not.

More permissions, and a real question about why. A visual companion may ask for the camera, the photo library, or the microphone. Some of those requests are inherent to a feature you want and some are not, and the distinction is worth making deliberately rather than at the moment the prompt appears. Camera access in particular is worth pausing over: an animated character does not need to see you, so if the request appears, the feature it serves should be identifiable. What each grant actually covers is in what a permission prompt actually grants.

A larger local footprint, and therefore more residue. Avatar assets, voice models and cached scenes take storage, and what is stored locally is what ends up in device backups and in whatever remains after you uninstall. That is a mild consideration rather than an alarming one, and it is covered in what ends up in your phone backup.

More upstream traffic, if voice is involved. Text is small. Audio is not, and if speech is transcribed on a server rather than on the device, what leaves your phone is a recording of your voice rather than a line of text. That is a materially different thing to have sent, and it is the one addition on this list that is worth checking before use rather than after.

What the word “virtual” tends to indicate in a listing

Mostly it indicates presentation, and it is worth being clear that presentation is not a capability.

An avatar makes an app feel more present without making its replies more coherent, more consistent, or better at retaining what you told it. The two properties are produced by unrelated parts of the system, and in practice a heavy visual layer sometimes accompanies a thinner conversational one, because rendering is where the effort went. That is not a rule and it is not a reason to avoid visual apps — it is a reason not to read the avatar as evidence about the conversation.

The vocabulary point generalises, and is set out at more length in what the words in these listings actually mean.

Why this post is short

Because padding it would mean inventing distinctions that do not exist. “Virtual AI friend app” and “AI friend app” describe the same category with the same evaluation problems, and the honest version of this page is a three-item supplement to the main guide rather than a second full treatment of the same material.

If you came here looking for how to judge one of these apps, the main guide is the page you want. If you came looking for whether the visual version needs different care, the answer is: a little, in the three specific ways above, and not otherwise.