Skip to content
AI Observatory · Italian tourism

Italy/Method

Method

How we measure, in order. Every choice here is a design decision, not a result: results are on the other pages.

01

In short

Every month we ask the same set of travel questions to AI engines with web search: ChatGPT, Perplexity, Copilot, Grok and Gemini.

Each question is asked 5 times on the same day on every engine. The 5 answers become a single value, so model randomness is not mistaken for change.

For every answer we look at two things: who is named (destinations, hotels) and who is the source (the links cited), split into DMO, Hotel, OTA and Other.

Every month rankings are compared with the previous month's, on shared questions only: the ± next to each number is that change.

02

Questions

The set covers seven categories: sea, mountains, lakes, villages, cities, snow, and the region in general (“what to do in Puglia”). Spas have been removed.

Two levels: comparing destinations (“where to go to the sea in Puglia”) and inside a destination (“best hotels in Monopoli”). Questions cover the whole trip (where to go, what to do, where to stay) and, for hotel questions, traveller types: families, couples, business, pets, budget, luxury. Phases are used to build a balanced set; on the site results add them up.

Every question is justified by a keyword with at least 1,000 searches a year in its language (Italian in Italy, English worldwide), measured on Google Ads through DataForSEO: volume of the exact keyword, sum of the last 12 months. It is the volume of that phrase (“where to stay costa smeralda”), not of the place name, which is searched far more. The keyword and its volume are under each question.

Destinations come from official sources: ISTAT tourist nights, Ministry lists, OpenSkiMap ski areas, OpenStreetMap parks, Most Beautiful Villages of Italy. Areas (parks, ski areas) are asked as “where to stay near X”.

The base is completed by the answers. Places AI proposes in one month that no official list has (Lake Bilancino, a minor ski area) become candidates for the following month, if at least two engines propose them in one answer out of four and they pass the same search threshold. A wave's set never changes while the wave is running.

The full list of questions, each with its keyword and yearly searches, can be downloaded as CSV below.

03

Engines

Answers are collected through Scrapeless, which queries the engines as a user would. Web search is forced where possible: ChatGPT and Perplexity with search on, Copilot in search mode, Grok in Expert mode.

Country: Italy for Italian questions; for English ones, Italy, except Gemini, which uses the United States because with Italy it answers in Italian.

The model each engine declares is stored with every answer: a model change shows in the data.

04

Waves and the 5 runs

A wave is a fixed day each month. The 5 runs start together, shuffled, so none systematically falls at a different moment.

A failed task is retried the same day. If an answer is still missing, its cell counts fewer answers: the value stays a fraction (k of n) and n is published.

Every answer is kept in full. The provider's raw payload is archived separately and is used to re-read an answer if the mapping changes.

05

Sources: DMO, Hotel, OTA, Other

A source is classified by its domain only, never by the text. Rules, in order: hand-written registry; OTA brands; municipality websites; the site of a hotel named in the same answer; everything else is Other.

Who tells the destination is measured without the hotel questions (“what are the best hotels in…”): there the question itself asks for hotels and OTAs, and they would weigh on every share.

DMO: tourism sites of public bodies (national, regions, municipalities, park authorities). A portal with an institutional-sounding name run by a private company is Other: the registry records the check.

OTA: online agencies and metasearch. Tripadvisor and Google Hotels are OTAs. Google cards cited by Gemini are OTA “Google Hotels” when they belong to a property, Other “Google Maps” otherwise.

Hotel: the official site of a property or a chain.

Other: guides, media, blogs, social, private portals, non-tourism institutions. The most cited domains get a label and a subtype.

06

Named hotels and destinations

Hotels: a language model (gpt-oss-120b, temperature 0) reads every answer and returns the named properties as JSON. Checks without a model: the name must appear verbatim in the text; a name with no property word that is the real name of a municipality, area or lake is discarded. The names are used to recognise hotels' official sites among the sources, and to show them by name instead of by address.

Places: on comparison questions the same model (temperature 0) reads the question and the answer and returns the places the answer proposes: lakes when lakes are asked for, ski resorts when the question is where to ski. Places named only as a reference ("a few kilometres from Florence") do not count. Checks without a model: the name must appear verbatim in the text; identity comes from the gazetteer (GeoNames Italy) when it knows the name, otherwise the place enters as a discovered one.

The extractor has a version, stored with every mention. A new version can be re-run on past waves without querying the engines again.

07

Metrics

Cell = one question on one engine, with its 5 same-day answers. For every entity (source type, DMO, hotel, place, region): presence = answers where it appears ÷ answers in the cell; first = answers where it is the first source or first named ÷ answers in the cell.

Every published number is the plain average of the cells in the selection: every question weighs the same, with no search-volume weights.

Change (±) = this month's presence minus the previous month's, in points, computed on questions present in both months only.

08

Rankings and movement

Sorted by presence, then by first. Equal values share a rank. No minimum threshold.

The path: a category in Italy gives the ranking of regions (only those where the category exists); a region gives the ranking of places and how much DMOs cover it; a place gives the ranking of hotels AI recommends.

DMOs and hotels have separate rankings. OTAs and other sources stay as a comparison.

The arrow is the rank movement versus the previous month: ▲ up, ▼ down, = stable, ● new.

09

Limits

An AI answer changes with the user, history, location and day: we measure an anonymous user, from a declared country, on one day.

A domain outside the registry and the rules stays Other: rarely cited DMOs can end up there until they are added to the registry.

The hotel extractor can be wrong: discarded mentions stay in the database with their reason.

A destination named negatively (“avoid X in August”) still counts as a mention.

10

Versions

Every wave records the version of the question set, the extractor and the source registry. A published wave is never recomputed: an error is fixed by republishing it with an erratum.

Questions

Every question of the wave with category, phase, place, keyword and yearly searches.

Download the questions (CSV) ↓