An AI Overview competitor monitor stores every cited domain for a fixed query set and compares those observations over time. Keep source citations separate from brand mentions, exclude failed requests from citation metrics, and publish the number of eligible runs beside every rate.
An AI Overview competitor monitor records every domain cited for a fixed list of Google searches. Run the same queries with the same country and language, store each source URL, then compare domains across dates. This shows who appeared in the responses you checked.
The monitor measures observations, not market share. A cited domain appeared in one captured result. The data cannot prove why Google selected that page, whether every user saw it, or whether your competitor will appear again tomorrow.
The Google AI Overviews API uses GET /ai-overview with query, gl, and hl. See the Google AI Overviews API product page for the current request plans.
Inspect the sources for one query
What should the competitor monitor measure?
Start with returned source records. In the live test from August 5, 2026, each source could contain url, title, site, and position. One source record in one dated response is the smallest useful observation.
| Value | Meaning | Keep separate from |
|---|---|---|
| Cited domain | A normalized hostname from a returned source | A brand name written in the answer |
| Cited URL | The exact page returned by the API | The domain used for grouping |
| Source position | The observed position value in that response | A classic Google ranking position |
| Brand mention | A company name found inside ai_overview | A linked citation |
A company can be mentioned without a link to its website. A domain can also be cited without the company name appearing in the answer. Combining both events into one score hides that difference.
Use the same test set for every competitor
Choose a query list related to the problems, comparisons, and products in your market. Tag the queries by topic if different groups have different importance. Then keep these inputs fixed:
- Query text and query ID
- Country in
gland language inhl - Collection schedule and timestamp format
- Rules for normalizing domains and counting subdomains
Do not create a separate collector for each competitor. Store every returned domain. A site you did not expect may become the most frequently cited source in the next period.
A small SQLite data model
The example monitor uses two tables. overview_run stores the request and its state. overview_citation stores the source rows connected with that request.
| Table | One row represents | Important fields |
|---|---|---|
overview_run | One query checked at one time | Query, country, language, state, answer, error |
overview_citation | One source returned for one run | Domain, URL, title, position |
This layout keeps failed requests even though they have no citation rows. That matters because a server error is not evidence that every competitor received zero citations.
A complete Python and SQLite monitor
The script creates the database, calls the API for every query, stores the response state, normalizes source domains, and prints a basic competitor report.
import os
import sqlite3
from datetime import datetime, timezone
from urllib.parse import urlparse
import requests
API_URL = "https://google-ai-overviews.p.rapidapi.com/ai-overview"
DATABASE_PATH = "ai_overview_competitors.db"
COUNTRY = "us"
LANGUAGE = "en"
QUERIES = [
"how does a heat pump work",
"heat pump maintenance checklist",
]
SCHEMA = """
CREATE TABLE IF NOT EXISTS overview_run (
id INTEGER PRIMARY KEY AUTOINCREMENT,
query_text TEXT NOT NULL,
country_code TEXT NOT NULL,
language_code TEXT NOT NULL,
observed_at TEXT NOT NULL,
request_state TEXT NOT NULL,
answer_text TEXT,
error_text TEXT
);
CREATE TABLE IF NOT EXISTS overview_citation (
id INTEGER PRIMARY KEY AUTOINCREMENT,
run_id INTEGER NOT NULL,
source_position INTEGER,
source_domain TEXT NOT NULL,
source_url TEXT NOT NULL,
source_title TEXT,
FOREIGN KEY (run_id) REFERENCES overview_run(id)
);
"""
def normalize_domain(source):
if not isinstance(source, dict):
return None
value = source.get("site") or source.get("url")
if not isinstance(value, str) or not value.strip():
return None
candidate = value.strip().lower()
if any(character.isspace() for character in candidate):
return None
if "://" not in candidate:
candidate = f"https://{candidate}"
host = urlparse(candidate).hostname
if not host:
return None
host = host.rstrip(".")
return host[4:] if host.startswith("www.") else host
def fetch_overview(query, api_key):
response = requests.get(
API_URL,
headers={
"x-rapidapi-host": "google-ai-overviews.p.rapidapi.com",
"x-rapidapi-key": api_key,
},
params={"query": query, "gl": COUNTRY, "hl": LANGUAGE},
timeout=45,
)
response.raise_for_status()
payload = response.json()
if not isinstance(payload, dict):
raise RuntimeError("The API response is not a JSON object.")
data = payload.get("data")
if payload.get("status") != 200 or not isinstance(data, dict):
message = payload.get("description") or "Unexpected API response"
raise RuntimeError(message)
answer_value = data.get("ai_overview")
answer = answer_value.strip() if isinstance(answer_value, str) else None
if not answer:
answer = None
raw_sources = data.get("sources")
if not isinstance(raw_sources, list):
raw_sources = []
citations = []
for source in raw_sources:
if not isinstance(source, dict) or not isinstance(source.get("url"), str):
continue
domain = normalize_domain(source)
if not domain:
continue
position = source.get("position")
if type(position) is not int:
position = None
citations.append(
{
"position": position,
"domain": domain,
"url": source["url"],
"title": source.get("title")
if isinstance(source.get("title"), str)
else None,
}
)
if not answer:
state = "no_overview"
elif not citations:
state = "no_sources"
else:
state = "success"
return state, answer, citations
def save_run(connection, query, state, answer=None, citations=None, error=None):
citations = citations or []
observed_at = datetime.now(timezone.utc).isoformat()
with connection:
cursor = connection.execute(
"""
INSERT INTO overview_run (
query_text, country_code, language_code, observed_at,
request_state, answer_text, error_text
) VALUES (?, ?, ?, ?, ?, ?, ?)
""",
(query, COUNTRY, LANGUAGE, observed_at, state, answer, error),
)
run_id = cursor.lastrowid
connection.executemany(
"""
INSERT INTO overview_citation (
run_id, source_position, source_domain, source_url, source_title
) VALUES (?, ?, ?, ?, ?)
""",
[
(
run_id,
citation["position"],
citation["domain"],
citation["url"],
citation["title"],
)
for citation in citations
],
)
def print_competitor_report(connection):
rows = connection.execute(
"""
SELECT
c.source_domain,
COUNT(DISTINCT c.run_id) AS cited_runs,
COUNT(*) AS source_slots,
MIN(c.source_position) AS best_observed_position
FROM overview_citation AS c
JOIN overview_run AS r ON r.id = c.run_id
WHERE r.request_state = 'success'
GROUP BY c.source_domain
ORDER BY cited_runs DESC, source_slots DESC, c.source_domain
"""
).fetchall()
print("\nCompetitor citation report")
for domain, cited_runs, source_slots, best_position in rows:
print(
f"{domain}: cited runs={cited_runs}, "
f"source slots={source_slots}, best position={best_position}"
)
def main():
api_key = os.environ.get("RAPIDAPI_KEY")
if not api_key:
raise RuntimeError("Set the RAPIDAPI_KEY environment variable first.")
connection = sqlite3.connect(DATABASE_PATH)
try:
connection.execute("PRAGMA foreign_keys = ON")
connection.executescript(SCHEMA)
for query in QUERIES:
try:
state, answer, citations = fetch_overview(query, api_key)
save_run(connection, query, state, answer, citations)
print(f"{query}: {state}, citations={len(citations)}")
except (requests.RequestException, RuntimeError) as error:
save_run(connection, query, "request_failed", error=str(error)[:500])
print(f"{query}: request_failed ({error})")
print_competitor_report(connection)
finally:
connection.close()
if __name__ == "__main__":
main()
Install the dependency with python -m pip install requests. Set RAPIDAPI_KEY, replace the sample queries, and run the script. It creates ai_overview_competitors.db in the current directory.
The report discovers domains from the response instead of using a fixed competitor list. Review the first results before deciding which domains belong to direct competitors, publishers, marketplaces, or unrelated sources.
Collect the first competitor baseline
Three metrics that remain understandable
| Metric | Formula | What it answers |
|---|---|---|
| Cited runs | Distinct successful runs containing the domain | How many observed responses cited it |
| Source slots | All source rows belonging to the domain | How many returned citation places it occupied |
| Best observed position | Lowest returned position value | The best position seen in the source array |
Do not combine these values into an unexplained visibility score. Publish the number of eligible runs beside every rate or share. Ten citations from 20 successful runs mean something different from ten citations collected from 2,000 runs.
Compare matching periods
Compare periods built from the same query IDs, country, language, and schedule. If the query set changed, compare only the intersection or label the periods as unmatched.
| Change | Evidence to keep |
|---|---|
| Newly cited domain | The first successful run and matching URL |
| Lost citation | The previous match and current successful non-match |
| Different cited page | The previous and current URLs for that domain |
| Brand mention changed | The two answer snapshots, analyzed separately from links |
A screenshot helps verify one response, but it cannot establish a trend. The database needs repeated observations under the same test conditions.
Handle failed and missing results correctly
Use at least four request states: success, no_sources, no_overview, and request_failed. Only successful runs with the source data required by a metric belong in that metric's denominator.
The August test produced temporary HTTP 500 responses. Counting one as zero citations would make every domain look weaker even though the collector failed. A valid request without an AI Overview is different again. Google says AI Overviews appear only when its systems decide they add value to classic Search.
What the monitor cannot prove
The dataset shows what the API observed for the supplied inputs. It does not explain why Google cited a page or prove that one competitor owns a topic. Google also documents that its AI features can use related searches across subtopics and that responses and links can vary.
Keep interpretations in a notes field. If a competitor page appears to cover a missing subtopic, record that as a hypothesis for editorial review, not as a fact returned by the API.
Use the single-domain citation checker for a quick site audit. Before mixing organic positions with this dataset, read why Google rankings and AI Overview citations can differ.
Frequently asked questions
Do I need a competitor list before collecting data?
No. Store every returned domain first. Classify the relevant domains after you inspect the data.
Is a brand mention the same as a citation?
No. A mention is text in the answer. A citation is a returned source record with a page URL. Keep separate columns for them.
Can I compare two different query lists?
You can, but the result is not a clean period comparison. Use the queries present in both periods or explain that the sample changed.
Sources and verification dates
- DataOcean Google AI Overviews on RapidAPI, listing checked September 22, 2026.
- Let's Scrape Google AI Overviews API product page, endpoint and request inputs checked September 22, 2026.
- Google Search Central documentation for AI features, checked September 22, 2026.
- Google's guide to generative AI features in Search, checked September 22, 2026.
The response fields and temporary HTTP 500 behavior in this guide come from a live product test performed on August 5, 2026. They are dated observations, not guarantees about future responses.