Supplementary Data Package (de-identified, truncated excerpts)
Manuscript: Competing Proximities: How Host-City Conditions Reorganise International News in Chinese Urban Press
Journal target: Chinese Journal of Communication

============================================================
CONTENTS
============================================================
1. Supplementary_Data_deidentified.csv
   - N = 1343 articles (analysis corpus)
   - Analysis variables + truncated text excerpts (max 400 characters)
   - 1015 / 1343 rows are truncated (excerpt_truncated = yes)
   - Full article texts are NOT included

2. reliability_v3_deidentified/
   - Inter-model reliability materials without full article republication
   - dual_llm_input_deidentified.csv: article_id, host_window, text_excerpt
     (max 500 characters; titles omitted)
   - Coded labels, alpha summaries, codebook, compute script retained

3. This README

Related narrative supplements (separate files in /supplementary):
- Supplementary_File_S1.md (robustness)
- Supplementary_File_S2_Codebook.md (coding tables)

============================================================
DE-IDENTIFICATION / TRUNCATION POLICY
============================================================
Removed:
- Full body_text (newspaper archive full texts are not redistributed)
- author (reporter bylines)
- Reliability raw API jsonl files that embed full article text
- Article titles in the reliability input file (reduce re-identification
  via headline search while retaining a usable excerpt)

Retained for analytic / review use:
- Coded variables used in the manuscript tables
- Outlet (source) and section labels
- Truncated text_excerpt for spot-checking source-line and topic coding

Truncation rule:
- Main corpus: first ~400 characters, preferring a nearby
  sentence boundary; append “…” when truncated
- Reliability sample: first ~500 characters, same rule
- excerpt_truncated marks whether truncation occurred

These excerpts are provided for peer-review verification of coding
decisions. They are not a licence to republish archive content.

============================================================
VARIABLE DICTIONARY — Supplementary_Data_deidentified.csv
============================================================
article_id         Stable numeric id matching the internal corpus index
host_window        2001_Shanghai | 2014_Beijing | 2026_Shenzhen | routine_other | unparsed
section            Newspaper section label (as archived)
page               Page marker (as archived; may be blank)
source             Outlet / masthead label within the Shenzhen Press Group ecology
date_text          Publication date string (as archived)
keyword_hits       Keyword hits used in retrieval (pipe-separated)
subject_roles      Coded subject-role labels (pipe-separated where multi-label)
event_topics       Coded event-topic labels (pipe-separated; multi-label)
story_form         Coded story-form category
source_line        Coded production/source-line category
sem_text_len       Approximate full-text length (characters); full text not shared
text_excerpt       Truncated article excerpt (max 400 chars)
excerpt_truncated  yes | no

============================================================
DATA AVAILABILITY NOTE FOR SUBMISSION PORTAL
============================================================
Upload this folder (or the zip) as Supplemental / Supporting files.
Select NO for “Is there a data set associated with this submission?”
unless/until a public repository DOI/URL is minted.

Suggested manuscript statement:
A de-identified analysis dataset with truncated article excerpts and
coding materials are provided as supplementary files with this submission.
Full article texts from the institutional newspaper archive are not
redistributed. Further materials are available from the corresponding
author upon reasonable request.
