README — SUPPLEMENTARY MATERIALS
==================================================================

Manuscript title: Whole-Exome Sequencing Reveals a Distinct Mutational
Landscape in Head and Neck Squamous Cell Carcinoma: A Pilot Study from
Northeast India

Journal: Molecular Biology Reports (Springer)
Corresponding author: Dr. Rajeev Kumar (rajeev.kvhfc@gmail.com)

This document lists and describes all supplementary files accompanying
the above manuscript. Two files are provided:

1. Supplementary_Methods.docx
2. Supplementary_Figures.docx
3. Supplementary_Tables.xlsx

------------------------------------------------------------------
FILE 1: Supplementary_Methods_and_Figures.docx
------------------------------------------------------------------

Contains:

  Supplementary Methods
    Detailed protocol for genomic DNA extraction, whole-exome
    sequencing (library preparation, capture, sequencing platform
    and run parameters), and the bioinformatics analysis pipeline,
    including read alignment, variant calling, VEP annotation
    parameters, and filtering thresholds. Also includes the detailed
    methodology for MAF file conversion and oncoplot generation
    (software, versions, and parameters used).

------------------------------------------------------------------
FILE 2: Supplementary_Figures.docx
------------------------------------------------------------------

Supplementary Figure S1 : Oncoplot of the top 30 recurrently mutated genes 
(post-variant prioritisation) in the combined cBioPortal CPTAC/TCGA-HNSCC dataset 
(GDC, 2025; n = 618)


------------------------------------------------------------------
FILE 3: Supplementary_Tables.xlsx
------------------------------------------------------------------

A single Excel workbook containing five tables, one per worksheet
tab. Tab names correspond to table numbers as listed below.

  Tab "Table S1" — Ensembl VEP-Annotated Somatic Variants
    Complete list of somatic variants identified by whole-exome
    sequencing across all ten HNSCC samples, annotated using Ensembl
    VEP (v115). Includes, at minimum: sample ID, chromosome,
    position, reference/alternate allele, gene symbol, transcript
    ID, consequence type (VEP IMPACT category), amino acid change,
    and any in silico deleteriousness scores (SIFT, PolyPhen-2,
    CADD, REVEL, AlphaMissense, SpliceAI) applied during filtering.

  Tab "Table S2" — ClinVar Pathogenic/Oncogenic Variants in
  Non-HNSCC-Associated Genes
    Full list of variants annotated as Pathogenic or Likely
    Pathogenic in ClinVar that map to genes without an established
    somatic role in HNSCC — i.e., genes with known Mendelian/
    germline disease associations rather than reported HNSCC driver
    or passenger status. Provided to support the manuscript's
    discussion of residual germline signal in a tumour-only calling
    approach. Includes gene symbol, variant (HGVS), ClinVar
    classification, associated Mendelian condition, and sample(s) in
    which the variant was detected.

  Tab "Table S3" — Per-Sample COSMIC SBS Signature Contributions
    Full quantitative breakdown of COSMIC Single Base Substitution
    (SBS) mutational signature contributions for each of the ten
    samples, reporting the relative proportion (%) attributed to
    each reference SBS signature (e.g., SBS1, SBS4, SBS5, SBS26,
    SBS86, etc.) per sample, underlying the signature summary
    presented in the main text and Supplementary Figure S3.

  Tab "Table S4" — Mutation Frequency Comparison Across Cohorts
    Mutation frequency of the top 200 OncoKB-annotated cancer genes,
    compared across three cohorts: the Northeast India (NEI) cohort
    (n=10, this study), TCGA-HNSC (n=618), and dbGENVOC (n=100).
    Includes gene symbol, OncoKB annotation tier, and per-cohort
    mutation frequency (% and sample counts) to allow direct
    cross-cohort comparison as referenced in the main text Results
    and Discussion. "Column Description provided besides the table"

Tab "Table S5" — Top 20 MODERATE/HIGH-impact variants of Tobacco-Exposure 
    Case–Control Variant Analysis, ranked by CADD score. 


------------------------------------------------------------------
GENERAL NOTES
------------------------------------------------------------------

- File formats: .docx (Microsoft Word, readable in Word 2016 or
  later, or any compatible word processor) and .xlsx (Microsoft
  Excel 2007 or later, or any compatible spreadsheet application).
- Abbreviations used across supplementary files are defined at
  first use in the main manuscript's abbreviations list; gene
  symbols follow HGNC nomenclature.
- Genome build: GRCh38/hg38 for all variant coordinates.
- For questions regarding any supplementary file, please contact
  the corresponding author at the address above.

==================================================================