An AI Overview competitor monitor records every domain cited for a fixed list of Google searches. Run the same queries with the same country and language, store each source URL, then compare domains across dates. This shows who appeared in the responses you checked.

The monitor measures observations, not market share. A cited domain appeared in one captured result. The data cannot prove why Google selected that page, whether every user saw it, or whether your competitor will appear again tomorrow.

The Google AI Overviews API uses GET /ai-overview with query, gl, and hl. See the Google AI Overviews API product page for the current request plans.

Inspect the sources for one query

What should the competitor monitor measure?

Start with returned source records. In the live test from August 5, 2026, each source could contain url, title, site, and position. One source record in one dated response is the smallest useful observation.

ValueMeaningKeep separate from
Cited domainA normalized hostname from a returned sourceA brand name written in the answer
Cited URLThe exact page returned by the APIThe domain used for grouping
Source positionThe observed position value in that responseA classic Google ranking position
Brand mentionA company name found inside ai_overviewA linked citation

A company can be mentioned without a link to its website. A domain can also be cited without the company name appearing in the answer. Combining both events into one score hides that difference.

Use the same test set for every competitor

Choose a query list related to the problems, comparisons, and products in your market. Tag the queries by topic if different groups have different importance. Then keep these inputs fixed:

  • Query text and query ID
  • Country in gl and language in hl
  • Collection schedule and timestamp format
  • Rules for normalizing domains and counting subdomains

Do not create a separate collector for each competitor. Store every returned domain. A site you did not expect may become the most frequently cited source in the next period.

A small SQLite data model

The example monitor uses two tables. overview_run stores the request and its state. overview_citation stores the source rows connected with that request.

TableOne row representsImportant fields
overview_runOne query checked at one timeQuery, country, language, state, answer, error
overview_citationOne source returned for one runDomain, URL, title, position

This layout keeps failed requests even though they have no citation rows. That matters because a server error is not evidence that every competitor received zero citations.

A complete Python and SQLite monitor

The script creates the database, calls the API for every query, stores the response state, normalizes source domains, and prints a basic competitor report.

import os
import sqlite3
from datetime import datetime, timezone
from urllib.parse import urlparse

import requests


API_URL = "https://google-ai-overviews.p.rapidapi.com/ai-overview"
DATABASE_PATH = "ai_overview_competitors.db"
COUNTRY = "us"
LANGUAGE = "en"
QUERIES = [
    "how does a heat pump work",
    "heat pump maintenance checklist",
]


SCHEMA = """
CREATE TABLE IF NOT EXISTS overview_run (
    id INTEGER PRIMARY KEY AUTOINCREMENT,
    query_text TEXT NOT NULL,
    country_code TEXT NOT NULL,
    language_code TEXT NOT NULL,
    observed_at TEXT NOT NULL,
    request_state TEXT NOT NULL,
    answer_text TEXT,
    error_text TEXT
);

CREATE TABLE IF NOT EXISTS overview_citation (
    id INTEGER PRIMARY KEY AUTOINCREMENT,
    run_id INTEGER NOT NULL,
    source_position INTEGER,
    source_domain TEXT NOT NULL,
    source_url TEXT NOT NULL,
    source_title TEXT,
    FOREIGN KEY (run_id) REFERENCES overview_run(id)
);
"""


def normalize_domain(source):
    if not isinstance(source, dict):
        return None

    value = source.get("site") or source.get("url")
    if not isinstance(value, str) or not value.strip():
        return None

    candidate = value.strip().lower()
    if any(character.isspace() for character in candidate):
        return None

    if "://" not in candidate:
        candidate = f"https://{candidate}"

    host = urlparse(candidate).hostname
    if not host:
        return None

    host = host.rstrip(".")
    return host[4:] if host.startswith("www.") else host


def fetch_overview(query, api_key):
    response = requests.get(
        API_URL,
        headers={
            "x-rapidapi-host": "google-ai-overviews.p.rapidapi.com",
            "x-rapidapi-key": api_key,
        },
        params={"query": query, "gl": COUNTRY, "hl": LANGUAGE},
        timeout=45,
    )
    response.raise_for_status()

    payload = response.json()
    if not isinstance(payload, dict):
        raise RuntimeError("The API response is not a JSON object.")

    data = payload.get("data")
    if payload.get("status") != 200 or not isinstance(data, dict):
        message = payload.get("description") or "Unexpected API response"
        raise RuntimeError(message)

    answer_value = data.get("ai_overview")
    answer = answer_value.strip() if isinstance(answer_value, str) else None
    if not answer:
        answer = None

    raw_sources = data.get("sources")
    if not isinstance(raw_sources, list):
        raw_sources = []

    citations = []
    for source in raw_sources:
        if not isinstance(source, dict) or not isinstance(source.get("url"), str):
            continue

        domain = normalize_domain(source)
        if not domain:
            continue

        position = source.get("position")
        if type(position) is not int:
            position = None

        citations.append(
            {
                "position": position,
                "domain": domain,
                "url": source["url"],
                "title": source.get("title")
                if isinstance(source.get("title"), str)
                else None,
            }
        )

    if not answer:
        state = "no_overview"
    elif not citations:
        state = "no_sources"
    else:
        state = "success"

    return state, answer, citations


def save_run(connection, query, state, answer=None, citations=None, error=None):
    citations = citations or []
    observed_at = datetime.now(timezone.utc).isoformat()

    with connection:
        cursor = connection.execute(
            """
            INSERT INTO overview_run (
                query_text, country_code, language_code, observed_at,
                request_state, answer_text, error_text
            ) VALUES (?, ?, ?, ?, ?, ?, ?)
            """,
            (query, COUNTRY, LANGUAGE, observed_at, state, answer, error),
        )
        run_id = cursor.lastrowid

        connection.executemany(
            """
            INSERT INTO overview_citation (
                run_id, source_position, source_domain, source_url, source_title
            ) VALUES (?, ?, ?, ?, ?)
            """,
            [
                (
                    run_id,
                    citation["position"],
                    citation["domain"],
                    citation["url"],
                    citation["title"],
                )
                for citation in citations
            ],
        )


def print_competitor_report(connection):
    rows = connection.execute(
        """
        SELECT
            c.source_domain,
            COUNT(DISTINCT c.run_id) AS cited_runs,
            COUNT(*) AS source_slots,
            MIN(c.source_position) AS best_observed_position
        FROM overview_citation AS c
        JOIN overview_run AS r ON r.id = c.run_id
        WHERE r.request_state = 'success'
        GROUP BY c.source_domain
        ORDER BY cited_runs DESC, source_slots DESC, c.source_domain
        """
    ).fetchall()

    print("\nCompetitor citation report")
    for domain, cited_runs, source_slots, best_position in rows:
        print(
            f"{domain}: cited runs={cited_runs}, "
            f"source slots={source_slots}, best position={best_position}"
        )


def main():
    api_key = os.environ.get("RAPIDAPI_KEY")
    if not api_key:
        raise RuntimeError("Set the RAPIDAPI_KEY environment variable first.")

    connection = sqlite3.connect(DATABASE_PATH)
    try:
        connection.execute("PRAGMA foreign_keys = ON")
        connection.executescript(SCHEMA)

        for query in QUERIES:
            try:
                state, answer, citations = fetch_overview(query, api_key)
                save_run(connection, query, state, answer, citations)
                print(f"{query}: {state}, citations={len(citations)}")
            except (requests.RequestException, RuntimeError) as error:
                save_run(connection, query, "request_failed", error=str(error)[:500])
                print(f"{query}: request_failed ({error})")

        print_competitor_report(connection)
    finally:
        connection.close()


if __name__ == "__main__":
    main()

Install the dependency with python -m pip install requests. Set RAPIDAPI_KEY, replace the sample queries, and run the script. It creates ai_overview_competitors.db in the current directory.

The report discovers domains from the response instead of using a fixed competitor list. Review the first results before deciding which domains belong to direct competitors, publishers, marketplaces, or unrelated sources.

Collect the first competitor baseline

Three metrics that remain understandable

MetricFormulaWhat it answers
Cited runsDistinct successful runs containing the domainHow many observed responses cited it
Source slotsAll source rows belonging to the domainHow many returned citation places it occupied
Best observed positionLowest returned position valueThe best position seen in the source array

Do not combine these values into an unexplained visibility score. Publish the number of eligible runs beside every rate or share. Ten citations from 20 successful runs mean something different from ten citations collected from 2,000 runs.

Compare matching periods

Compare periods built from the same query IDs, country, language, and schedule. If the query set changed, compare only the intersection or label the periods as unmatched.

ChangeEvidence to keep
Newly cited domainThe first successful run and matching URL
Lost citationThe previous match and current successful non-match
Different cited pageThe previous and current URLs for that domain
Brand mention changedThe two answer snapshots, analyzed separately from links

A screenshot helps verify one response, but it cannot establish a trend. The database needs repeated observations under the same test conditions.

Handle failed and missing results correctly

Use at least four request states: success, no_sources, no_overview, and request_failed. Only successful runs with the source data required by a metric belong in that metric's denominator.

The August test produced temporary HTTP 500 responses. Counting one as zero citations would make every domain look weaker even though the collector failed. A valid request without an AI Overview is different again. Google says AI Overviews appear only when its systems decide they add value to classic Search.

What the monitor cannot prove

The dataset shows what the API observed for the supplied inputs. It does not explain why Google cited a page or prove that one competitor owns a topic. Google also documents that its AI features can use related searches across subtopics and that responses and links can vary.

Keep interpretations in a notes field. If a competitor page appears to cover a missing subtopic, record that as a hypothesis for editorial review, not as a fact returned by the API.

Use the single-domain citation checker for a quick site audit. Before mixing organic positions with this dataset, read why Google rankings and AI Overview citations can differ.

Frequently asked questions

Do I need a competitor list before collecting data?

No. Store every returned domain first. Classify the relevant domains after you inspect the data.

Is a brand mention the same as a citation?

No. A mention is text in the answer. A citation is a returned source record with a page URL. Keep separate columns for them.

Can I compare two different query lists?

You can, but the result is not a clean period comparison. Use the queries present in both periods or explain that the sample changed.

Sources and verification dates

The response fields and temporary HTTP 500 behavior in this guide come from a live product test performed on August 5, 2026. They are dated observations, not guarantees about future responses.

Start a dated competitor citation baseline