Searching your own brand name is the most common mistake. If you ask ChatGPT "what is <your brand>" it will usually say something — that only tells you the model has heard of you. What decides whether you get considered is the buyer-intent question: "best <category> tool for <situation>", "alternatives to <competitor>", "which <category> should I choose for <constraint>". Your visibility is whether you appear in THOSE answers.
The second mistake is the single check. Ask the same question ten times and the answer set changes: different framing, different brands, sometimes different sources. A single run is one draw from a distribution, so a yes/no from one query carries almost no information. The honest unit of measurement is a rate over repeated samples — and the width of the interval around that rate tells you whether a change is real or just noise.
The third thing to separate is being MENTIONED, being CITED and being RECOMMENDED. An engine can name you in passing, link you as a source, or actually tell the user to choose you. Those are different outcomes with different commercial value, and they need to be measured separately rather than collapsed into one 'visibility' number.
Finally, ChatGPT is not one surface. Answers with live browsing/citations behave differently from answers generated purely from the model's parametric memory, and the same question can resolve differently in each mode. A measurement that does not say which surface it observed is ambiguous by construction.