CONFIDENTIAL CASE FILE
📁 Return to the
Front
Evidence Source
Exhibit A: Breast Cancer Dataset
This evidence file draws on the Breast Cancer dataset curated by
Reihaneh Namdari and hosted on Kaggle.
It is a cleaned subset of the US Surveillance, Epidemiology, and End
Results (SEER) cancer registry and focuses on women diagnosed with
infiltrating duct and lobular carcinoma between 2006 and 2010.
-
Source: SEER population-based cancer registry (US
National Cancer Institute), curated for Kaggle.
-
Sample size: 4,024 female patients with breast cancer.
-
Time frame: Diagnoses from 2006–2010, with follow-up
for survival in months.
-
Kaggle Dataset Page
Key Evidence Fields
Each “subject” in this case file includes 14 variables capturing
demographic, tumor, and outcome information:
-
Demographics: Age at diagnosis, race, marital status.
-
Tumor characteristics: Tumor size, histologic
grade/differentiation, T stage, N stage, overall cancer stage, and
extent of regional spread.
-
Lymph node findings: Number of regional lymph nodes
examined and number found positive.
-
Hormone receptor status: Estrogen receptor (ER) and
progesterone receptor (PR) status.
-
Outcomes: Survival time in months and vital status
(alive vs deceased).
How This Evidence Will Be Used
In this investigation, the dataset serves as the primary evidence for
exploring how clinical and pathological features of breast cancer relate
to patient survival. Using statistical techniques, we will:
-
Profile the “suspect” tumors by stage, receptor status, and lymph node
involvement.
-
Identify patterns associated with better or worse survival outcomes.
-
Build predictive models that estimate prognosis based on baseline
clinical characteristics.
Think of this dataset as the master case ledger: each
row is a patient, each column is a clue, and our analytic tools will
determine whether the evidence is strong enough to convict breast cancer
of its deadly impact.