How to Evaluate an AI Friend App Before You Pay for One

Evaluating one of these apps works best as a sequence rather than a checklist, because the checks have wildly different costs. Two minutes on the store listing eliminates more candidates than an hour of use, the first session tells you almost everything about conversational quality, and the questions that actually decide whether you keep paying only become answerable after a few weeks. Running the stages in the wrong order means spending money to learn things that were free.

What follows is that sequence. It assumes you cannot trust marketing copy, cannot trust published rankings, and are not going to install six apps to compare them.

Stage one: before you install

Everything here is free, takes about two minutes, and rules out the majority of listings.

Check who publishes it. Tap the developer name and look at their other apps. A coherent portfolio or nothing else is normal; a spread across unrelated genres is the strongest negative signal a store listing contains. The full set of provenance checks is in spotting a lookalike companion app.

Open the linked privacy policy in a browser before installing. You are not reading it closely yet. You are checking that the domain matches the app’s claimed identity, that the document is about this app rather than a template with someone else’s product name left in, and that it has a last-updated date.

Read the declared data categories in the store listing and compare them against what the app says it does. A text chat app declaring precise location or contacts is declaring capability its features do not need.

Read the one-star reviews only. The overall star average is nearly useless; the low-rated reviews are where billing failures, cancellation problems and sudden behaviour changes get reported. You are looking for repeated specifics, not sentiment.

Check whether an export exists, usually mentioned in the listing or the policy. Whether you can get your conversation history out determines how much a future switch costs you, and it is the single most consequential feature nobody looks for.

Stage two: the first session

Install, do not pay, and spend twenty minutes deliberately rather than casually.

Establish the character and then test whether it holds. Talk for ten minutes, then change subject abruptly, then come back. Voice drift within one session is common and it tells you how thin the standing instructions are.

Plant one specific, unremarkable detail early — a stated preference, a name, a plan for Thursday. Do not mention it again. This is the test you will use later.

Watch what happens at the limits. Say something the app is likely to refuse. The interesting thing is not whether it refuses but how: a clean acknowledgement is a designed boundary, whereas an abrupt topic change, a generic deflection, or a reply that is visibly a different voice is a filter operating on top of the model. Which of those you are seeing matters for predictability.

Turn the network off and try to send a message. Ten seconds, and it settles where the generation happens.

Note where the paywall sits. Whether it appears before any usable functionality, after a fixed number of messages, or at emotionally-loaded moments is a design decision that tells you what the product optimises for.

Stage three: before you pay

This is the stage most people skip, and it is the one that costs money to skip.

Find the cancellation path first. Not the price — the cancellation. If you are paying through a platform store, the store’s cancellation screen works whether or not the app cooperates. If you are paying an operator directly by card, you are relying entirely on that operator’s own process. The difference is set out in who you actually bought the subscription from, and it matters more than the amount.

Establish what the paid tier actually changes. Every product in this category prices something different — volume, features, the character’s behaviour, or the relationship framing itself. Whatever it is, check it is a thing you noticed wanting during stage two rather than a thing the paywall told you to want.

Now read the policy properly, specifically the sections on retention, on whether conversations are reviewed by people, on use for model training, and on deletion. Read what the document says rather than what a review says it says. If any of those four is absent, that absence is the finding.

Decide what you are prepared to have stored. This is a decision about your own conduct rather than about the app, and it is the only privacy control that works regardless of what the operator does.

Stage four: after a few weeks

Three questions that cannot be answered earlier and decide whether the thing is worth keeping.

Did the planted detail survive? Ask about it now. This is the only honest test of a memory claim, and the result is usually less impressive than the marketing. Why that is structural rather than a failing of any one product is covered in how AI companion apps work.

Has the behaviour changed under you? Operators alter models, instructions and filters without notice. Note whether the app you are paying for is still the app you evaluated.

Is your own usage what you expected it to be? Not a moral question — a practical one. If the pattern of use has drifted a long way from what you signed up for, that is worth noticing regardless of what conclusion you draw.

What this sequence will not settle

None of these stages tells you whether an app is good. They tell you whether it is what it claims, whether you can leave, and whether its own claims survive contact with use. Conversational quality is a matter of fit between a product and a person, and it does not generalise — which is why no ranking of these apps can be trusted and why this post recommends nothing.

The sequence also cannot verify anything about what happens on the operator’s side. Retention, human review, training use and third-party sharing are matters of published commitment, not observation, and no amount of testing from the outside reaches them. That asymmetry is permanent, and the correct response is to decide what you are willing to send rather than to hope for a way to check.