# Codebook: ranked lists and publisher self-ranking (paper 3, Study A) Version 1.1 (2026-09-27). Any change after annotation starts is logged in the "Revision log" at the end, with the items affected and whether they were re-coded. ## Your task Each row of your sheet (`sheet_A.csv` or `sheet_B.csv`) is one web page, as it was fetched on 26–27 September 2026. Open the saved copy in `data/processed/listicles/annotation_pages/.html`. Use the live URL only if the saved copy is unreadable, and write "used live page" in `notes` when you do (the live page may have changed). Code every item independently. Do not discuss items with the other annotator and do not look at their sheet. You will not see what the automatic detector decided; that is deliberate. ## Fields **is_ranked_list** — `yes` / `no` / `unsure` `yes` if the main content presents three or more named products, companies, services or places as a list of recommendations ("best X", "top N", "our picks"), whether numbered or not. `no` for articles, product pages, category pages, directories without editorial selection, and lists of tips or steps. If `no`, leave the remaining fields blank except `publisher_name`, `disclosure` and `notes`. **first_five_entries** — the first five recommended entities, in page order, separated by ` ; `. Use the entity's name only ("Rippling", not "1. Rippling: best for payroll"). Ignore navigation, sidebars, sponsor boxes, "related articles" and tables of contents. If the page has a quick summary list and then detailed sections, use the order of the **main detailed sections**, and note "summary order differs" if it differs. **n_entries** — how many recommended entities the main list has (count them; approximate above 30 as `30+`). **publisher_name** — the organisation that publishes the page, from page evidence: logo, header, footer copyright, "about us". Do **not** infer it from the web address alone. If the page is on a marketplace or blogging platform (e.g. medium.com), the publisher is the platform unless the page is clearly a company's own blog there. **publisher_in_list** — `yes` if the publisher, or a product or service it sells under its own brand, is one of the list's entries. Sister brands of the same company count as `yes`; write the relationship in `notes`. **publisher_position** — the publisher's position in the main list (1 = first). Blank if not in the list. **order_is_ranking** — `yes` if the page presents its order as a ranking (numbers, "#1", "best overall" first); `no` if it says the order is alphabetical or not a ranking; `unsure` otherwise. **disclosure** — `yes` if the page tells the reader that the publisher sells one of the listed products or has a commercial relationship with listed companies (e.g. "we are the makers of X", "we may earn a commission"). `no` otherwise. **notes** — anything that made the item hard: summary vs detailed order, several lists on one page, the page is not in English, the saved copy is broken, etc. ## Edge cases - Several lists on one page: code the first list that answers the page's headline. - A "#1 pick" box above the list that repeats an entry: ignore the box; code the list. - A sponsored or advertisement slot inside the list: code it as an entry and say so in `notes`. - Product variants of the same brand ("Acme Pro", "Acme Lite"): each variant is an entry. - Displayed rank numbers win over page order. In a countdown (#20 shown first, #1 last), `first_five_entries` starts with the entity shown as #1, then #2, and so on, and `publisher_position` is the displayed rank. Without displayed numbers, use page order. Write "countdown" in `notes`. ## Second task: mention sheet (`mention_sheet.csv`, 120 rows) Each row pairs an entity name with one AI answer (`mention_answers/.md`). Read the whole answer. **mentioned** — `yes` if the answer names this entity in any form (abbreviation, product of the same brand, spelling variant); `no` otherwise. **recommended_first** — `yes` if it is the first entity the answer recommends; `no` otherwise; blank if not mentioned. Note unusual forms in `notes` (e.g. "named only as 'the Acme app'"). ## Revision log - v1.1 (2026-09-27, after model annotator A batch 1 of 12, before any other batch): added the countdown rule (displayed rank numbers win over page order). Batch 1 was re-checked under v1.1 by the same annotator. - Correction (2026-09-27, after coding, before any validity statistic): model annotator B's batch 2 wrote five rows under the wrong item ids (a permutation: content for L282 under L043 and vice versa; L065 under L047, L176 under L065, L047 under L176). Each row's notes describe the page it was coded from, and a name-overlap check against every page confirmed the permutation. The rows were moved to their correct ids; the original file is kept as `model_B/batch_02.original_misaligned.csv.bak`. No other batch failed the check.