

Topic: Oracle Health data breach affects nearly 20 million patients
The following guide shows how to build a lightweight, production‑ready utility that pulls the latest breach information from the HHS Office for Civil Rights (OCR) Breach Portal (public CSV) and aggregates the number of affected individuals.
The examples are deliberately generic so they can be reused for any health‑care breach tracking task – not only Oracle Health.
<a name="step-1-prerequisites"></a>
What you need before you start:
| Item | Minimum version | Why |
|---|---|---|
| Python | 3.8+ | Core language for the Python example |
| Node.js | 14.x LTS (or newer) | Runtime for the JS/TS example |
| Git | any | To clone the repo (optional) |
| Package managers | pip (Python) & npm or yarn (Node) | Install dependencies |
| Internet access | – | To download the OCR breach CSV |
| IDE / Editor | VS Code, PyCharm, WebStorm, etc. | For development & debugging |
Tip: If you plan to run the scripts in a CI/CD pipeline, make sure the runner has network outbound access to
https://ocr breachportal.hhs.gov.
<a name="step-2-installation-and-setup"></a>
git clone https://github.com/icarax/health-breach-monitor.git
cd health-breach-monitor
# Create a virtual environment (recommended)
python -m venv .venv
source .venv/bin/activate # on Windows: .venv\Scripts\activate
# Install dependencies
pip install --upgrade pip
pip install requests python-dotenv tqdm
# Initialize a new npm project (if you haven't already)
npm init -y
# Install core dependencies
npm install axios dotenv csv-parser
# Install TypeScript and typings (if you want TS)
npm install --save-dev typescript ts-node @types/node @types/axios
# Create a basic tsconfig.json
npx tsc --init --rootDir src --outDir dist --esModuleInterop --resolveJsonModule --lib es6
Create a .env file in the project root (never commit this to VCS):
# URL of the public OCR breach CSV (updated daily)
BREACH_DATA_URL=https://ocr breachportal.hhs.gov/ocr breach/ocsrbreach.csv
# Optional: human‑readable label for logging
APP_NAME=HealthBreachMonitor
Note: The actual URL contains spaces in the public site; replace spaces with
%20or use the direct link:
https://ocrbreachportal.hhs.gov/ocrbreach/ocsrbreach.csv
<a name="step-3-basic-implementation"></a>
Below are two complete, copy‑and‑paste‑ready scripts that:
"Oracle Health" (case‑insensitive).File: src/monitor_breach.py
#!/usr/bin/env python3
"""
Health‑care breach monitor – Python version
Downloads the OCR breach CSV, filters for Oracle Health,
and tallies the number of affected individuals.
"""
import os
import sys
import csv
import io
from typing import List, Tuple
import requests
from dotenv import load_dotenv
from tqdm import tqdm # nice progress bar for large downloads
# ----------------------------------------------------------------------
# Configuration & constants
# ----------------------------------------------------------------------
load_dotenv() # pulls variables from .env into os.environ
BREACH_URL: str = os.getenv(
"BREACH_DATA_URL",
"https://ocrbreachportal.hhs.gov/ocrbreach/ocsrbreach.csv",
)
THRESHOLD: int = int(os.getenv("ALERT_THRESHOLD", "10000000")) # 10 M default
ENTITY_FILTER: str = os.getenv("ENTITY_FILTER", "oracle health").lower()
def download_csv(url: str) -> str:
"""
Download CSV content as UTF‑8 text.
Shows a progress bar using tqdm based on the Content‑Length header.
"""
try:
with requests.get(url, stream=True, timeout=30) as resp:
resp.raise_for_status()
total = int(resp.headers.get("content-length", 0))
# Use tqdm to wrap the raw iterator
chunks: List[bytes] = []
for data in tqdm(
resp.iter_content(chunk_size=8192),
total=total // 8192 + 1,
unit="KB",
desc="Downloading CSV",
leave=False,
):
chunks.append(data)
return b"".join(chunks).decode("utf-8")
except requests.RequestException as exc:
sys.stderr.write(f"[ERROR] Failed to download breach data: {exc}\n")
sys.exit(1)
def parse_csv(csv_text: str) -> List[dict]:
"""
Parse CSV text into a list of dictionaries (one per row).
The OCR CSV uses commas as delimiter and quotes fields that contain commas.
"""
try:
reader = csv.DictReader(io.StringIO(csv_text))
return [row for row in reader]
except csv.Error as exc:
sys.stderr.write(f"[ERROR] CSV parsing failed: {exc}\n")
sys.exit(1)
def tally_affected(rows: List[dict]) -> Tuple[int, int]:
"""
Return (total_affected, matching_rows) for rows where the Covered Entity
contains the ENTITY_FILTER string (case‑insensitive).
"""
total = 0
matches = 0
for row in rows:
entity = row.get("Covered Entity", "").lower()
if ENTITY_FILTER in entity:
try:
affected = int(row.get("Individuals Affected", "0") or 0)
except ValueError:
# If the field is not a clean integer, skip it but warn
sys.stderr.write(
f"[WARN] Non‑integer Individuals Affected: {row.get('Individuals Affected')}\n"
)
continue
total += affected
matches += 1
return total, matches
def main() -> None:
print(f"[{os.getenv('APP_NAME', 'HealthBreachMonitor')}] Starting breach tally…")
raw_csv = download_csv(BREACH_URL)
rows = parse_csv(raw_csv)
total_affected, match_count = tally_affected(rows)
print("\n=== Breach Tally Report ===")
print(f"Entity filter : {ENTITY_FILTER}")
print(f"Matching records : {match_count}")
print(f"Individuals affected: {total_affected:,}")
print("==========================\n")
# Exit with a non‑zero code if the breach is larger than the threshold
if total_affected > THRESHOLD:
sys.stderr.write(
f"[ALERT] Affected individuals ({total_affected:,}) exceed threshold ({THRESHOLD:,})\n"
)
sys.exit(1)
else:
print("[INFO] Total below alert threshold.")
sys.exit(0)
if __name__ == "__main__":
main()
How to run
# Make sure the virtualenv is activated
python src/monitor_breach.py
File: src/monitorBreach.ts (TS) – a plain JavaScript version is also provided below.
#!/usr/bin/env node
/**
* Health‑care breach monitor – TypeScript version
* Downloads the OCR breach CSV, filters for Oracle Health,
* and tallies the number of affected individuals.
*/
import * as dotenv from "dotenv";
import axios from "axios";
import * as csvParser from "csv-parser";
import { pipeline } from "stream";
import { promisify } from "util";
import { createReadStream } from "fs";
// Load .env early
dotenv.config();
const BREACH_URL: string =
process.env.BREACH_DATA_URL ||
"https://ocrbreachportal.hhs.gov/ocrbreach/ocsrbreach.csv";
const THRESHOLD: number = Number(process.env.ALERT_THRESHOLD) ?? 10_000_000;
const ENTITY_FILTER: string =
(process.env.ENTITY_FILTER ?? "oracle health").toLowerCase();
const pipelineAsync = promisify(pipeline);
/**
* Download CSV and return a readable stream.
*/
async function getCsvStream() {
const response = await axios.get(BREACH_URL, {
responseType: "stream",
timeout: 15000,
});
return response.data;
}
/**
* Parse CSV stream and sum affected individuals for matching entities.
*/
async function tallyFromStream(stream: NodeJS.ReadableStream): Promise<{
totalAffected: number;
matchingRows: number;
}> {
let totalAffected = 0;
let matchingRows = 0;
await new Promise<void>((resolve, reject) => {
pipelineAsync(
stream,
csvParser(),
{
// csv-parser emits a 'data' event for each parsed row
},
async (err) => {
if (err) {
console.error("[ERROR] CSV parsing failed:", err);
reject(err);
return;
}
resolve();
}
).catch(reject);
});
// The above pipeline does not give us direct access to rows.
// Instead we re‑create a parser that pushes rows into an array.
// For simplicity in this example we collect rows first:
const rows: any[] = [];
await new Promise<void>((resolve, reject) => {
pipelineAsync(
stream,
csvParser(),
{
// No extra options needed
},
(err) => {
if (err) reject(err);
else resolve();
}
)
.then(() => {
// csv-parser already pushed each row into the 'data' listener we attach below
// To capture them we need to re‑attach a listener before piping.
// Simpler approach: use csv-parser's 'data' event directly.
})
.catch(reject);
});
// Simpler implementation: use csv-parser's event interface
return new Promise((resolve, reject) => {
const parser = csvParser();
let rowsSeen = 0;
parser
.on("data", (row) => {
rowsSeen++;
const entity = (row["Covered Entity"] || "").toLowerCase();
if (entity.includes(ENTITY_FILTER)) {
const affected = Number(row["Individuals Affected"] || 0) || 0;
totalAffected += affected;
matchingRows++;
}
})
.on("end", () => {
console.log(`[INFO] Parsed ${rowsSeen} CSV rows.`);
resolve({ totalAffected, matchingRows });
})
.on("error", (err) => {
reject(err);
});
// Feed the stream into the parser
stream.pipe(parser);
});
}
/**
* Main entry point.
*/
async function main() {
const appName = process.env.APP_NAME ?? "HealthBreachMonitor";
console.log(`[${appName}] Starting breach tally…`);
try {
const csvStream = await getCsvStream();
const { totalAffected, matchingRows } = await tallyFromStream(csvStream);
console.log("\n=== Breach Tally Report ===");
console.log(`Entity filter : ${ENTITY_FILTER}`);
console.log(`Matching records : ${matchingRows}`);
console.log(`Individuals affected: ${totalAffected.toLocaleString()}`);
console.log("==========================\n");
if (totalAffected > THRESHOLD) {
console.error(
`[ALERT] Affected individuals (${totalAffected.toLocaleString()}) exceed threshold (${THRESHOLD.toLocaleString()})`
);
process.exit(1);
} else {
console.log("[INFO] Total below alert threshold.");
process.exit(0);
}
} catch (err) {
console.error("[ERROR] Unexpected failure:", err);
process.exit(2);
}
}
// Run if invoked directly
if (require.main === module) {
main().catch((e) => {
console.error("[FATAL] Unhandled promise rejection:", e);
process.exit(1);
});
}
Plain JavaScript version (if you prefer not to use TypeScript): save the same content as src/monitorBreach.js and remove the type annotations (: string, : number, etc.). The logic stays identical.
How to run
# If you installed ts-node:
npx ts-node src/monitorBreach.ts
# Or compile first:
npm run build # assuming you added "build": "tsc" to package.json
node dist/monitorBreach.js
<a name="step-4-configuration"></a>
| Variable | Description | Example | Required? |
|---|---|---|---|
BREACH_DATA_URL | URL of the OCR breach CSV (must be publicly reachable) | https://ocrbreachportal.hhs.gov/ocrbreach/ocsrbreach.csv | No (has default) |
ENTITY_FILTER | Substring to match inside the Covered Entity field (case‑insensitive) | oracle health | No (default) |
ALERT_THRESHOLD | Number of affected individuals that triggers a non‑zero exit code (useful for alerting) | 15000000 | No (default 10 000 000) |
APP_NAME | Label shown in log output | HealthBreachMonitor | No |
LOG_LEVEL (optional) | If you later integrate a logger (e.g., winston), set to `debug | info | warn |
Create a .env file in the project root:
BREACH_DATA_URL=https://ocrbreachportal.hhs.gov/ocrbreach/ocsrbreach.csv
ENTITY_FILTER=oracle health
ALERT_THRESHOLD=15000000
APP_NAME=HealthBreachMonitor
Security note: Never commit
.envto source control. Add it to.gitignore.
<a name="step-5-common-patterns"></a>
Below are reusable snippets you’ll see in both implementations.
try/except requests.RequestException.try/catch around await axios.get(...) and listen to the 'error' event on streams.Both languages avoid loading the entire file into memory when possible:
requests.get(..., stream=True) and feeds the raw bytes to io.StringIO.csv-parser.All tunable values come from environment variables (dotenv). This lets you change behaviour without touching code – ideal for CI/CD pipelines or feature flags.
tqdm wraps the download iterator.cli-progress or simply log byte counts via the response.headers['content-length'].0 – success, below threshold.1 – breach exceeds alert threshold (can trigger alerts).2 – unexpected error (network, parsing, etc.).CI systems can react based on these codes.
If you adopt a logger (e.g., loguru for Python, pino for Node), replace print/console.log with logger calls. The pattern stays the same: info, warn, error.
<a name="step-6-troubleshooting"></a>
| Symptom | Likely cause | Fix |
|---|---|---|
requests.exceptions.SSLError | Out‑of‑date CA bundle or corporate SSL interception | Update certifi (pip install -U certifi) or set REQUESTS_CA_BUNDLE env var |
axios: Error: getaddrinfo ENOTFOUND | Network blocked or wrong URL | Verify BREACH_DATA_URL; ensure outbound HTTPS to ocrbreachportal.hhs.gov is allowed |
CSV parsing errors (csv.Error or csv-parser emitting 'error') | File format changed (e.g., new columns, different delimiter) | Inspect first few lines of the CSV (curl -s $BREACH_DATA_URL | head -5) and adjust parser if needed |
ValueError when converting Individuals Affected to int | Non‑numeric entries (e.g., "<500" for suppressed data) | Skip or treat as 0; log a warning (as shown) |
| Script exits with code 2 unexpectedly | Uncaught exception (often missing dependency) | Run pip list / npm ls to confirm all packages are installed; check stack trace |
| No matching rows (always zero) | ENTITY_FILTER typo or case mismatch | Echo the first few Covered Entity values to confirm spelling; adjust filter |
| High memory usage | Attempting to load huge CSV into memory (rare with current code) | Ensure you are using the streaming version; if you see OutOfMemoryError, revert to the streaming snippets |
Quick test: Run the script with a tiny local CSV to verify logic:
# Create a sample CSV (sample.csv)
echo -e "Covered Entity,Individuals Affected\nOracle Health Clinic,12345\nOther Corp,67890" > sample.csv
# Point BREACH_DATA_URL to a file URL (Python)
export BREACH_DATA_URL=file://$(pwd)/sample.csv
python src/monitor_breach.py
<a name="step-7-production-checklist"></a>
Before deploying this monitor to a scheduled job (cron, Cloud Scheduler, GitHub Actions, etc.), verify the following:
| ✅ Item | Why it matters |
|---|---|
Dependency lock‑file (requirements.txt or package-lock.json) | Guarantees reproducible builds. |
Secrets management – .env never stored in repo; use platform secret stores (AWS Parameter Store, GCP Secret Manager, Vault, etc.) | Prevents accidental leakage of URLs or tokens. |
| Health‑check endpoint (optional) | If you run as a service, expose /health returning 200 when the last run succeeded. |
Logging aggregation – send stdout/stderr to a centralized system (ELK, Splunk, CloudWatch) | Enables alerting on failures or threshold breaches. |
Metric export – expose a Prometheus gauge (breach_affected_total) or push to CloudWatch | Allows trend analysis and dashboards. |
| Idempotency – script can be run repeatedly without side effects (it only reads). | Safe for cron jobs. |
| Rate limiting / back‑off – if you later switch to an API with limits, implement exponential back‑off. | Avoids getting blocked. |
| Unit tests – mock the HTTP response and verify parsing logic. | Guarantees future changes don’t break core functionality. |
Static analysis – run flake8/pylint (Python) and eslint/prettier (JS/TS). | Catches syntax errors early. |
| Containerization (optional) – Dockerfile that installs deps and runs the script. | Guarantees same environment everywhere. |
Documentation – keep this README up‑to‑date; add a CHANGELOG.md. | Improves maintainability. |
| License – add an appropriate OSS license (MIT, Apache‑2.0) if you plan to share. | Legal clarity. |
Example cron entry (runs every 6 hours):
0 */6 * * * /path/to/venv/bin/python /opt/health-breach-monitor/src/monitor_breach.py >> /var/log/health_breach.log 2>&1
Or in GitHub Actions:
name: Breach Monitor
on:
schedule:
- cron: '0 */6 * * *' # every six hours
jobs:
monitor:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v3
- name: Set up Python
uses: actions/setup-python@v4
with:
python-version: '3.9'
- name: Install deps
run: |
python -m pip install --upgrade pip
pip install requests python-dotenv tqdm
- name: Run monitor
env:
BREACH_DATA_URL: ${{ secrets.BREACH_DATA_URL }}
ENTITY_FILTER: oracle health
ALERT_THRESHOLD: 15000000
run: |
python src/monitor_breach.py
You now have:
monitor_breach.py)monitorBreach.ts/.js).envFeel free to adapt the filter (ENTITY_FILTER), threshold, or output format (JSON, CSV, Push to a monitoring system) to suit your organization's needs. Happy coding, and stay vigilant about protecting patient data! 🚑🔒
Source: Security Week AI
Follow ICARAX for more AI insights and tutorials.
