1. Welcome to RangerIO

RangerIO Is Your AI Safety Gateway

RangerIO sits on your laptop between your sensitive data and the AI tools you already use. Files come in — contracts, ledgers, EHR exports, research datasets. RangerIO scans them for sensitive fields, profiles them into a structured dossier, and caches that understanding locally. When you want cloud AI in the loop, RangerIO masks the sensitive values before they leave the workstation and restores them when the answer comes back.

   Your files                                        Your AI tools
   ──────────       ┌─────────────────────────┐      ───────────────
   DOCX, XLSX  ───▶ │   Privacy Scan          │ ───▶ Claude
   PDF, DB     ───▶ │   Safe Handoff          │ ───▶ ChatGPT
   API         ───▶ │   Local Memory          │ ───▶ Perplexity
                    └─────────────────────────┘
                       RangerIO Gateway
                       (runs on your laptop)

The result: you can finally use Claude, ChatGPT, and Perplexity on the work you couldn't share with them before — without exposing the underlying data, and without re-paying to rebuild the same understanding every time.

The Three Pillars

Privacy Scan. Sensitive fields are detected on ingestion and tiered by confidence — confirmed, suspected, dismissed. You curate what's safe to share. Your overrides stick.

Safe Handoff. When cloud AI is in the loop, RangerIO masks sensitive values before sending and restores them when the answer returns. The cloud sees stand-ins, never the real data.

Local Memory. Every dossier stays on your laptop and gets read by everything else — chat, reports, the connector. RangerIO never charges you to re-derive what it already knows about your file.

Who RangerIO Is For

  • Compliance and engagement professionals at SMB scale — auditors, healthcare administrators, legal and HR ops, accounting firms, IoT and industrial compliance.
  • Individual professionals working with sensitive client material on their own laptop.
  • Buyers who want to use the AI tools they already pay for on the work they couldn't share with them before.

RangerIO is not a self-serve developer tool, not a hyperscaler ML platform, and not a SaaS data lake. See What RangerIO Doesn't Do for the honest boundary.

Two Editions

EditionWhat's BundledBest For
FullIncludes a local LLM runtime so RangerIO can answer locally without any external model.A laptop with a discrete GPU or Apple Silicon, or anyone who wants RangerIO to be fully self-contained.
LiteSame data engine and dossier — connect your own local or cloud model when you want answers.A standard 16 GB business laptop, or anyone who already pays for Claude / ChatGPT / a self-hosted model.

The data engine works the same in both. Intelligence (LLM-driven analysis) is layered on top and upgradeable later — you can start in Lite and add a model when ready.

What's Next


2. Why RangerIO Exists

The Friction

You know the pause. The second before pasting a contract into ChatGPT. The second-guess before uploading a client list. That isn't paranoia — it's the part of your job AI can't help with yet.

Cloud AI works — until your data is the input. The moment the work involves a confidential agreement, a patient record, an audit file, or anything else with real sensitivity, the most useful tool on your desk turns into a liability. So you carve out the AI work into a smaller, less interesting pile, and do the real work the way you've always done it. By hand.

Three Things Wrong With Cloud-Only AI

Ungrounded. A cloud model has no idea what's in your file. It answers from generic patterns it learned from public data, not from your data. That's the structural cause of hallucination — not a tuning problem, a grounding problem.

Exposed. When you paste sensitive content into a cloud chat, it crosses your perimeter. Vendor terms vary, retention policies change, and an honest mistake (an autocomplete pulling in a regulated field) is enough to be a compliance event.

Expensive. You're already paying for Claude, ChatGPT, or Perplexity. You'd like to use them on more of your work, not pay a second AI vendor to redo it locally.

What RangerIO Solves

Edge-first means RangerIO does the heavy understanding on your laptop. Files get profiled into a dossier — schema, entities, quality issues, sensitive fields — that's cached locally and read by every other capability. The model you use most often is the local cache, not a cloud round-trip.

Cloud-as-escalation means RangerIO doesn't pretend the cloud has nothing to offer. When you want Claude or ChatGPT in the loop, RangerIO is the gateway that gets your data ready: sensitive values masked, dossier shipped instead of raw rows, answers restored when they come back.

The combination — local profiling + governed cloud handoff — is what turns the AI subscription you already pay for into something you can finally use on the work that matters most.

Why Now

Three forces converged:

  1. Local AI got good enough. Small models running on consumer hardware are now strong enough to do meaningful profiling, classification, and chat. The compute math changed.
  2. Connector standards landed. The MCP protocol (and the family of standards alongside it) means a single gateway surface can talk to Claude Desktop, ChatGPT, Perplexity, and whatever ships next year — without you re-integrating each time.
  3. Compliance fatigue hit a wall. SMB-scale regulated professionals can't wait two years for hyperscaler-grade tools to fit their budget and risk profile. They need the laptop they already have to be the deployment unit.

RangerIO is what falls out when you take all three seriously.


3. How RangerIO Works

The Gateway in Three Steps

When you bring a file into RangerIO, three things happen — in order, automatically:

Step 1 — Privacy Scan

Every column and every text field is checked for sensitive content: names, addresses, contact details, government IDs, financial identifiers, dates that re-identify, locations, organizations. Findings are tiered by confidence (confirmed / suspected / dismissed) so you can see at a glance what's certain vs. what needs a second look. A sanity gate filters out structural false positives — for example, financial dollar columns that an over-aggressive matcher would otherwise flag as credit cards.

Step 2 — Profile Into a Dossier

While the Privacy Scan runs, RangerIO is also building the dossier — a structured understanding of the file. Schema, types, column statistics, quality issues, entities, and the PII findings from Step 1 all land in the dossier. It's the canonical thing every other RangerIO capability reads from: chat, reports, the cloud connector.

Step 3 — Cache Locally

The dossier stays on your laptop, in Local Memory. The next question against the same file doesn't re-do the profile work. The next month's similar engagement starts with overrides and curation already in place. The cloud connector reads from the dossier, not from the raw rows.

What Happens When You Use Cloud AI

If you stop at Step 3, RangerIO is a private workbench. If you go further and connect a cloud AI tool (Stage 2 — Claude Desktop today; ChatGPT and Perplexity follow), the Safe Handoff kicks in:

   You ask a question in Claude Desktop
                     │
                     ▼
   RangerIO masks sensitive values  ◀── Privacy Scan tells it what to mask
                     │
                     ▼
   Stand-ins go to the cloud, not real data
                     │
                     ▼
   The model answers using the masked context
                     │
                     ▼
   RangerIO restores real values in the answer
                     │
                     ▼
   You see your real data; the cloud never did

The salt that links real values to stand-ins lives only on your laptop.

Where Things Live

ThingWhere it livesWhat the cloud sees
Your raw filesOn your laptopNever
The dossierOn your laptopThe summary, by your choice
Sensitive valuesOn your laptopStand-ins only
The mapping from stand-in back to real valueOn your laptopNever
Your AI tool subscriptionCloudYour masked queries

Hardware Tiers, in Plain Terms

RangerIO classifies your hardware on first launch and adapts. You don't configure this.

TierWhat you haveWhat it gets you
Tier 1 — Lite Standard16 GB RAM, no GPUFull data engine; bring your own AI
Tier 2 — Lite Plus16-32 GB, modest GPU or strong CPULite SKU with low-latency profiling
Tier 3 — Full Standard16-32 GB, mid-range discrete GPU or Apple Silicon 16 GB+Full SKU with a small bundled model
Tier 4 — Full Plus32-64 GB, larger discrete GPUFull SKU with a larger model
Tier 5 — Workstation64 GB+, high-end GPU or Apple Silicon 32 GB+Multiple models loaded; large datasets

If RangerIO can't fit a model in available memory at startup, you'll see a prompt with three options: retry after closing apps, fall back to a smaller model, or skip the AI layer for now. The data engine never blocks on model availability.


4. System Requirements & Installation

System Requirements

Supported Platforms

  • Windows 10 / 11 (64-bit)
  • macOS 12 (Monterey) or later — Apple Silicon (M1+) and Intel both supported

Minimum Hardware (Lite SKU)

  • 16 GB RAM
  • Modern multi-core CPU (Intel/AMD x86-64 or Apple Silicon)
  • 8 GB free disk space (more if you plan to add models)
  • No GPU required

Recommended for Full SKU

  • 32 GB RAM
  • Discrete GPU (NVIDIA 8 GB+ VRAM) or Apple Silicon 16 GB+ unified memory
  • 50 GB free disk space (room for multiple models and project data)

See the tier table on How RangerIO Works for what each hardware level unlocks.

Disk Space Breakdown

  • Base install: ~5 GB
  • Local model (Full SKU): 2-8 GB depending on the model bundled
  • Project data: scales with your imports — plan ~3× the size of your largest dataset for working space

Installation

Windows

  1. Download RangerIO-Setup.exe from the website.
  2. Double-click to run.
  3. Default install path: C:\Program Files\RangerIO. Change it if you prefer.
  4. Launch from the Start Menu when install completes.

On first launch, Windows Defender may scan the application — this is normal. The Windows Firewall will ask permission for local network access; allow it (RangerIO uses localhost for its web interface).

macOS

  1. Download RangerIO.dmg from the website.
  2. Open the DMG and drag RangerIO to Applications.
  3. First launch: right-click the app and choose Open (Gatekeeper requires this once).
  4. Accept the security prompt.

Apple Silicon Macs use a native arm64 binary.

First Launch

When you first open RangerIO:

  1. Hardware classification. RangerIO measures your CPU, RAM, GPU, and VRAM and picks a tier.
  2. Resource readiness check. If memory is tight, you'll see a prompt offering: retry after closing apps, fall back to a smaller model, or skip the AI layer for now.
  3. Model selection (Full SKU). Choose a bundled model that fits your tier, or skip and use Lite mode.
  4. Data location. Pick where projects and dossiers are stored. Default: ~/RangerIO/ on Mac, %LOCALAPPDATA%\RangerIO\ on Windows.
  5. Ready. RangerIO opens to the project list.

Lite Mode (No Local Model)

Lite mode uses the data engine without bundling a model. You can still:

  • Import files and build full dossiers
  • See the Privacy Scan, quality profile, and column intelligence
  • Curate PII findings
  • Export Markdown dossiers

When you're ready for AI-assisted answers, attach a local model (any GGUF-format llama.cpp model on disk) or connect to your existing AI tool through the Stage 2 connector.

Updating

RangerIO does not auto-check for updates. To update: download the new installer from the website, close RangerIO, and run it. Your projects, dossiers, models, and settings are preserved.

Uninstalling

Windows: Settings → Apps → RangerIO → Uninstall. You'll be asked whether to keep your project data.

macOS: Drag RangerIO from Applications to the Trash. Project data lives at ~/RangerIO/ and ~/Library/Application Support/RangerIO/ — remove these manually if you want a clean uninstall.


5. Bringing Data In

Projects

A project is a workspace. It holds:

  • One or more imported files or database connections
  • The dossier RangerIO built for each
  • Your PII overrides
  • Chat history and reports

Use one project per engagement, client, or analysis. Rename or archive when done.

To create a project, click New Project on the home screen, give it a descriptive name (Q4_Vendor_Audit reads better than data1 two months later), and click Create.

Supported Formats

Tabular Files

  • CSV, TSV (any delimiter; auto-detected)
  • Excel (.xlsx, .xls) — multi-sheet workbooks profiled per sheet, with a workbook-level summary
  • JSON, JSONL (nested objects supported)
  • Parquet

Documents

  • PDF (text and layout extracted)
  • Microsoft Word (.docx)
  • Plain text and Markdown

Databases

  • PostgreSQL
  • MySQL / MariaDB
  • Microsoft SQL Server
  • SQLite

APIs

  • REST endpoints with optional bearer-token authentication

Import Methods

Drag-and-Drop

The fastest way in. Drag files from your file manager onto the project window and drop on the import area. RangerIO detects the format, infers types, profiles the file into a dossier, runs the Privacy Scan, and marks the source ready.

Multiple files and folders are supported in a single drop.

File Upload (Detailed)

Click Import Data → Upload Files when you need more control over parsing. You can configure:

  • CSV/TSV: delimiter, header row, encoding, date format
  • Excel: which sheets to import, header row offset
  • JSON: root path for nested data, array vs. object mode

Database Connection

Click Import Data → Connect to Database. Provide host, port, database name, and credentials. Read-only accounts are recommended.

Test the connection before importing. Then choose:

  • Import Table — copy data into RangerIO's local store. Best when the data is stable.
  • Live Query — query directly against the source on every question. Best when the data changes frequently.

You can also configure sampling: import all rows, first N rows, or a random sample.

Directory Watching

Point RangerIO at a folder and it will import new files as they appear. Use file-glob filters (sales_*.xlsx, *.csv) to scope what gets pulled in. Choose a one-time scan or continuous monitoring.

Common use cases: automated daily reports, log analysis, batch processing pipelines.

API Connection

Click Import Data → API Connection. Provide:

  • URL — the endpoint
  • Method — GET (most common)
  • Headers — include Authorization: Bearer <token> if required
  • JSON path — the path inside the response that holds the data array
  • Refresh schedule — manual, or every N hours/days

RangerIO sends a test request before saving the connection.

What You See After Import

Every source gets a dossier. The dossier card shows:

  • File name and source type
  • Row and column counts
  • Schema (column names + inferred types)
  • Quality summary
  • Privacy Scan summary
  • Status: ready, degraded, or blocked

Click into the source for full column statistics, PII findings, and the data preview.

Excel Workbooks

Excel files are profiled at two levels:

  1. Each sheet gets its own dossier.
  2. The workbook gets a parent dossier that aggregates findings across sheets — combined quality summary, combined PII findings per domain, combined entity profile.

You can ask questions of a single sheet or the whole workbook.

Refreshing a Source

To re-import a file or rerun a database query:

  1. Open Data Sources.
  2. Click the refresh icon on the source.
  3. RangerIO re-profiles the data and tells you what changed: row deltas, schema shifts, quality drift.

PII overrides you've set on a column persist across rescans. Curation gets sharper over time, not reset.


6. The Dossier

What It Is

The dossier is RangerIO's structured understanding of your file. It's built once at ingestion, cached locally, and read by every other capability — chat, reports, the Stage 2 connector — so you never re-pay to re-derive what RangerIO already knows. This local cache is what RangerIO calls Local Memory: every dossier you build adds to it, and every later question reads from it.

What's In a Dossier

When you click into a source you'll see:

  • Overview. Row and column counts, size, ingestion time, status.
  • Schema. Column names, inferred types, type-distribution sparklines.
  • Column Statistics. Per column: min, max, mean, median, distribution, missing-value rate, unique count, most common values.
  • Quality Profile. Format and encoding issues, whitespace anomalies, case inconsistencies, categorical drift. (RangerIO's quality checks focus on normalization issues — RangerIO does not flag missing values or "outliers" as quality problems.)
  • Entities. Names, locations, organizations, dates, and other entities detected in your data — both in structured columns and in narrative text columns.
  • PII Findings. Sensitive fields, tiered by confidence. See Privacy Scan.
  • Domain Extensions. When the source is recognized as belonging to a domain (financial, general), domain-specific intelligence is layered in — for example, a Benford's-law check on numeric columns in a financial source.
  • Status. ready, degraded, or blocked — RangerIO declares what it knows and what it's missing rather than guessing.

When the Dossier Updates

The dossier is rebuilt:

  • On import.
  • When you click Refresh on a source.
  • When you change PII overrides (only the affected sections rebuild).
  • After a model upgrade, on demand — you can rescan an older dossier to upgrade it from the heuristic tier to a model-backed tier without re-importing the data.

Workbook Aggregation

For Excel multi-sheet workbooks, the parent dossier aggregates child sheet findings:

  • Quality issues are summed.
  • Entity findings are unioned across sheets.
  • PII findings are merged with per-sheet attribution preserved.
  • The dossier preserves findings from every domain present in the workbook (finance, HR, healthcare, etc.). Reports today render the dominant domain by default; cross-domain rendering is on the roadmap.

This means a single question on a workbook reads from one combined dossier — without losing the sheet-level breakdown when you want it.

Markdown Export

You can export the full dossier as Markdown for sharing or archive: see Reports & Exports. Sensitive values are not included by default — you choose what travels.


7. Privacy Scan

What the Privacy Scan Does

When you import a file, RangerIO scans every column for sensitive information. Findings are tiered by confidence:

TierMeaning
ConfirmedValidator-corroborated and high-confidence — definitely sensitive
CriticalHigh-impact category (credit cards, government IDs)
SuspectedLikely sensitive but worth a human read
HeuristicPattern-based suggestion
DismissedAuto-rejected by the sanity gate
User-criticalA finding you marked as critical

The sanity gate is what makes the scan usable in production: it drops findings that lack structural corroboration so you don't drown in false positives. Without it, a financial column of dollar amounts would surface as 80+ phantom credit-card matches.

Curating Findings

You curate findings two ways.

From the Data Preview

Click the colored dot on a column header to open a popover. From there you can:

  • Mark a column as PII if RangerIO missed it.
  • Dismiss a false positive.
  • Override the type or severity — for example, downgrade an aggressive credit_card flag to account_number.
  • Add a reason so you remember why later.

From the Privacy Findings Panel

Open the Privacy Findings section in the side panel for a list view of every finding across the source. Same actions, batch-friendly.

Your Overrides Stick

Overrides persist across rescans. When you re-import a source or upgrade the dossier:

  • Findings you confirmed stay confirmed.
  • Findings you dismissed stay dismissed.
  • New findings RangerIO discovers show up as suspected for review.

This is the loop that makes the Privacy Scan get sharper over time on engagements you repeat.

What's Detected

RangerIO uses a single combined entity-extraction pass that covers the most common categories of sensitive information across structured columns and narrative text:

  • Names, addresses, contact details
  • Government IDs (SSN, SIN, etc. — varies by locale)
  • Financial identifiers (account numbers, credit cards)
  • Dates of birth and other dates that re-identify
  • Locations
  • Organizations and counterparties

For specialized identifiers like medical record numbers or contract reference codes, mark the column as PII manually from the column dot — your override carries the same persistence behavior as auto-detected findings.

A Single Source of Truth

The classifier output is the single ground truth. Other parts of RangerIO — compliance reports, the Stage 2 connector, exports — read from the classifier cache. There is no parallel scanning path that could disagree.

Compliance reports only count findings in compliance-grade tiers (confirmed, critical, user_critical). If a finding is dismissed or only heuristic, it does not appear in the compliance count.


8. Safe Handoff

What Safe Handoff Does

Safe Handoff is the pillar that runs whenever cloud AI is in the loop. It's how RangerIO turns "this data is too sensitive for ChatGPT" into "the cloud sees enough to help, never enough to leak."

How the Round Trip Works

   Your data on your laptop
            │
            ▼
   ┌──────────────────────────┐
   │ 1. Mask                  │   Sensitive values replaced with stable
   │    (Privacy Scan tells   │   stand-ins. The structure is preserved
   │    Safe Handoff what to  │   so the model can still reason about
   │    mask)                 │   shape, relationships, counts.
   └──────────────────────────┘
            │
            ▼
   ┌──────────────────────────┐
   │ 2. Send                  │   Masked content goes to the cloud
   │                          │   AI tool over your authorized channel.
   └──────────────────────────┘
            │
            ▼
   ┌──────────────────────────┐
   │ 3. Reason                │   The model answers using the masked
   │                          │   context. It never sees real values.
   └──────────────────────────┘
            │
            ▼
   ┌──────────────────────────┐
   │ 4. Restore               │   When the answer comes back, RangerIO
   │                          │   substitutes the real values back in.
   │                          │   You see your data; the cloud didn't.
   └──────────────────────────┘

Stable Stand-Ins

The stand-ins Safe Handoff produces are stable — the same real value gets the same stand-in throughout an engagement. That matters because:

  • The model can reason about repeated entities ("this same vendor appears in three contracts") without seeing the vendor name.
  • Joins across files still work — if acct_4099-2231 appears in your AP and in your bank reconciliation, the model can correlate them through the matching stand-in.
  • Round-trip restoration always lands on the right value.

Per-Engagement Salt

The mapping from real values to stand-ins is keyed by an engagement-scoped salt:

  • The salt rotates on engagement close, or every 90 days, whichever is sooner.
  • A registry maps stand-ins back to real values for round-trip restoration.
  • A 90-day hot replay window keeps the full registry available; after that, only the audit log is retained.

You don't manage the salt — RangerIO handles rotation. Engagement boundaries are set per project.

What the Cloud Never Sees

  • Raw rows from your data
  • Real sensitive values (names, IDs, account numbers, etc.)
  • The mapping from stand-in back to real value (it lives only on your laptop)
  • Cached intelligence with zero-state defaults that would falsely broadcast a "ready" signal (a defensive guard sanitizes this)

What the Cloud Does See

  • The masked content of your prompt
  • Schema-level information (column names, types) when you choose to share it
  • The dossier summary when the Stage 2 connector exposes it

You decide what travels. The default posture is conservative — share less than you think you need, expand from there.


9. Asking Questions

How Questions Work

Open the chat from the Assistant tab. Pick a source (or multiple) from the dropdown and ask in plain English.

RangerIO routes questions through a cascade of handlers, fastest to most expensive:

  1. Direct lookup — counts, schema queries, simple aggregates served from the dossier without invoking a model.
  2. Smart SQL — RangerIO translates the question to SQL and runs it on the local analytics engine when the question maps cleanly to a query.
  3. Grounded advisory — for analytical or interpretive questions, RangerIO assembles a grounded prompt: row count, column list, dossier summary, and (when relevant) a sample of rows. The model answers with the dossier's facts as the ground truth.

You don't pick the path — RangerIO picks based on the question. Each path returns the answer plus enough context to verify it.

What "Grounded" Means

Answers are anchored to your file's actual statistics and dossier — not to whatever the model thinks the world looks like. If you ask "how many rows?" and the dossier says 21, you get 21. The model is not allowed to invent a different number.

When the answer comes from Smart SQL, you'll see the SQL that ran. When it comes from grounded advisory, you'll see the dossier facts that anchored it.

Example Questions

Direct lookup

  • "How many rows are in this dataset?"
  • "What columns are in the Q4 vendor file?"
  • "Which columns have PII?"

Smart SQL

  • "Top 10 customers by revenue."
  • "All transactions over $10,000 from last quarter."
  • "Distinct product categories with their counts."

Grounded advisory

  • "Summarize the key trends in the sales data."
  • "Are there any unusual patterns in vendor billing?"
  • "Compare the quality profile of the Q3 and Q4 files."

Multiple Sources

Select multiple sources from the dropdown. RangerIO reads from each dossier and attributes citations back to the source that contributed each fact.

For Excel workbooks, you can ask the workbook (which uses the aggregated parent dossier) or a specific sheet.

When the Model Doesn't Have an Answer

If the question is interpretive and the dossier doesn't contain enough context, RangerIO returns a clear "I don't have enough to answer that" — not a fabricated number. This is the grounding contract applied to chat.

If you see this, the next step is usually:

  • Ask a more specific question.
  • Run a Privacy Scan re-curate so RangerIO has more confirmed structure.
  • Add a related source so cross-source context is available.

Performance

Direct lookups return in well under a second. Smart SQL runs at analytics-engine speed (typically sub-second on local files). Grounded advisory speed depends on the model — local Full SKU answers in a few seconds; Lite SKU with a connected cloud model is bounded by the network round-trip.


10. Connecting to Your AI Tools

The Stage 2 Connector

RangerIO works as a governed connector for the AI tools you already use. Today that's Claude Desktop. ChatGPT and Perplexity ship on the same surface in named phases.

The connector exposes RangerIO's dossier — not the raw rows. Your AI tool can:

  • Read each source's dossier, schema, and quality profile.
  • Run ingestion prompts (import a file, import a directory).
  • Read the Privacy Scan summary.
  • Read project-level summaries.

What the connector does not expose:

  • Raw rows from your data.
  • Sensitive values inside columns flagged as PII (Safe Handoff masks them first).
  • Cached intelligence with zero-state defaults that would broadcast a "ready" signal when no scan has actually run.

Setting Up Claude Desktop

The first time you enable the connector:

  1. From RangerIO, go to Settings → Connectors → Claude Desktop.
  2. Follow the on-screen instructions to add a RangerIO entry to Claude Desktop's claude_desktop_config.json. The location depends on your OS — RangerIO shows you the exact path and the snippet to paste.
  3. Restart Claude Desktop.
  4. The RangerIO connector appears in Claude Desktop's resource menu.

You'll see your projects, sources, and ingestion prompts there. Click any source to read its dossier in Claude.

The configuration is one-time per machine. Updates to the connector layer don't require re-configuration unless we explicitly note it in a release.

What "Vendor-Neutral" Means Here

The connector layer is intentionally generic — it does not encode anything specific to Claude Desktop. ChatGPT and Perplexity use the same surface; only the client-side configuration step differs. This is enforced in our build process: a check fails if any client-specific brand names leak into the connector code.

Round-Trip Masking

When the AI tool needs to reason about sensitive content (a contract clause, a clinical note paragraph), the Safe Handoff described in Page 8 kicks in. Masked content travels; restored answers come back.

Roadmap

  • Stage 2.0 — Claude Desktop (current).
  • Stage 2.1 — ChatGPT.
  • Stage 2.2 — Perplexity.

These ship on the same connector surface. No new install or migration step is required when later phases land.


11. Reports & Exports

Two Kinds of Output

OutputUse For
Markdown Dossier ExportSharing the structured intelligence about a source — schema, quality, entities, PII summary.
Structured ReportA specific report contract (e.g., a compliance summary, a quality summary) that declares what fields it needs and renders accordingly.

Markdown Dossier Export

From any source, click Export → Markdown. RangerIO emits a structured Markdown document containing:

  • Source identity (name, format, ingestion time, row/column counts)
  • Schema with inferred types
  • Column statistics
  • Quality summary
  • Entity findings (counts and categories)
  • PII summary (counts per tier; values are not included unless you explicitly opt in for a project)
  • Domain extensions, when present
  • Workbook breakdown (for Excel), when present

You can configure what's included before export. Sensitive values are redacted by default.

Structured Reports

Each report type declares the dossier fields it needs. When you generate a report, RangerIO checks the dossier and returns one of three states:

  • Ready — every required field is present; the report renders with full content.
  • Degraded — some required fields are present; the report renders with the available sections, marking the missing ones explicitly.
  • Blocked — too few required fields are present; RangerIO refuses to render rather than fabricate.

Reports never silently substitute hallucinated content for missing data. If you see a blocked report, the right next step is usually to re-curate the Privacy Scan or rescan with a stronger model.

Exporting Cleaned Data

When you want to export a working dataset (rather than a dossier or report):

  1. Open the source.
  2. Apply normalization actions if needed (whitespace, case, encoding fixes).
  3. Click Export → Data.
  4. Choose the format: CSV, Excel, JSON, JSONL, Parquet.
  5. Optionally apply a PII action: leave as-is, mask in place, or replace with stable stand-ins.

For most regulated workflows, mask in place is the right default — your column structure is preserved and downstream tools see the masked values.

Audit Trail

Every export records what was exported (source, columns, row count), what PII action was applied, and when it happened. The audit trail is local. It is included in the project archive if you back up the project.


12. Use Cases

Four common scenarios, each anchored to a buyer profile and the friction RangerIO removes.

Individual Professionals

The setup. A solo consultant, contractor, or domain expert working from their own laptop, juggling a mix of sensitive client material — research datasets, draft analyses, briefing memos.

Without RangerIO: the AI subscription you already pay for is the most useful thing on your desk for the work you can share, and useless for the work you can't. You manually carve client material into "OK to paste" and "definitely not."

With RangerIO:

HeadacheRangerIO
Sensitive files you can't paste into ChatGPT or ClaudeProfile them locally — then bring just the findings, not the raw rows, to the AI you already use
Data quality review needed, no analyst on handAutomated profiling in minutes — schema, gaps, anomalies, structure
Hundreds of pages, no time to read them allAsk in plain English; answers grounded in your actual data
PII tools that flag everything and miss what mattersTiered detection you can curate; overrides stick across rescans

Accounting & Audit

The setup. Engagement teams at SMB-scale firms — auditors, controllers, compliance leads — handling AP ledgers, journals, vendor contracts, and reconciliation files at month-end and quarter-end.

Without RangerIO: client books are too sensitive to paste into ChatGPT. Manual review takes the time you don't have. New AI features in Excel and Word can't see the data either.

With RangerIO:

HeadacheRangerIO
Books too sensitive for ChatGPTProfile locally, then take the dossier to your AI — never the raw rows
Duplicates, missing approvalsFlagged at ingestion, before your first question
PII across filesTiered detection you can curate — your overrides stick
New client every monthRe-run the dossier flow on each new file in minutes

Healthcare

The setup. Clinical operations leads, healthcare administrators, and research teams working with EHR exports, lab CSVs, billing files, and research datasets containing PHI.

Without RangerIO: PHI exposure rules out cloud AI for most real questions. De-identification by hand is slow and error-prone. Outcomes-research and quality-improvement work that would benefit from AI doesn't get the help.

With RangerIO:

HeadacheRangerIO
MRNs, DOBs, names scattered across filesTiered detection at the column level — you curate confirmed, suspected, dismissed
EHR exports, lab CSVs, billing files — every format differentOne ingestion flow turns each into the same dossier shape
IRB wants PHI scrubbed before analysis startsFindings ready before your first question; overrides stick across rescans
Outcomes research needs cloud AI; charts can't leaveSend the dossier — cohort counts, distributions, outliers — not the chart

Legal & HR

The setup. Counsel, paralegals, HR leads, and people-operations teams handling contracts, NDAs, case files, employee records, and policy documents.

Without RangerIO: privileged content can't go to ChatGPT, but the volume is too high for manual review. Counsel needs the analysis without the raw file. Employee data with SINs, salaries, and termination grounds is off-limits to most AI tools.

With RangerIO:

HeadacheRangerIO
SINs, salaries, termination grounds across HR filesTiered detection per column — you curate what's safe to share
Hundreds of pages of contracts and case files to reviewProfile each into a dossier — entities, dates, sensitive fields surfaced
Privileged content can't go to ChatGPT or ClaudeSend the dossier — not the document — to the AI you already use
Counsel needs analysis without the raw fileExport the dossier; originals stay on your laptop

13. What RangerIO Doesn't Do

Honest boundaries — so you can tell whether RangerIO fits your work, and so you can argue against the idea internally if it doesn't.

RangerIO Is Not...

A SaaS data lake or warehouse

RangerIO runs on your laptop. It is not a cloud platform you log into, and it does not store your data on someone else's infrastructure. The deployment unit is the workstation. If your team needs centralized storage with multi-user access controls and a dashboard for the CIO, you need a different tool.

A regulatory certification

RangerIO's architecture is designed for regulated workflows — local-first processing, governed cloud handoff, opt-in connectivity, no telemetry. But "HIPAA-ready", "SOC 2-certified", "PCI-DSS-compliant" are certifications that belong to your organization and your audit posture, not to a piece of software. RangerIO is part of your compliance program; it does not substitute for one.

A DLP or endpoint security tool

RangerIO does not monitor your filesystem for unauthorized data movement, scan emails leaving your network, or enforce policy at the OS level. It's a workbench for sensitive data, not a perimeter control. If you need DLP, you need a DLP — RangerIO complements it, doesn't replace it.

A self-serve developer tool

RangerIO has APIs and a vendor-neutral connector surface, and developers can build against it. But the product is shaped for compliance and engagement professionals, not for an engineer integrating it into a microservices stack. There is no SDK marketplace today and no public API for arbitrary extension.

A streaming or real-time pipeline

RangerIO ingests files and database snapshots. It is not Kafka, it is not a CDC pipeline, and it does not subscribe to event streams. If your data shows up as new rows in a database every minute, the Live Query mode against a database connection is your closest fit; if you need sub-second event ingestion with stream semantics, RangerIO is the wrong shape.

A replacement for your AI subscription

RangerIO does not replace Claude, ChatGPT, or Perplexity. It makes them usable on more of your work. If you don't already pay for or use a cloud AI, RangerIO's Full SKU bundles a local model so you have something to ask questions of — but the gateway story is most powerful when you're plugging in an AI you already trust.

Stage 3 yet

Stage 3 — the Enterprise Platform with skills marketplace and control plane — is on the multi-year roadmap. Today, RangerIO is Stage 1 (Workbench, shipping) and rolling out Stage 2 (Governed AI Layer — Claude Desktop now; ChatGPT and Perplexity to follow). If your evaluation depends on Stage 3 capabilities being shipped, you should wait.

Cloud-only, ever

RangerIO is edge-first. Even with Stage 2 connectors fully wired, the dossier, the salt, the registry that maps real values to stand-ins, and your raw data all stay on your laptop. If you need a true cloud-native AI platform with hyperscaler integration, RangerIO is the wrong shape.

What This Boundary Buys You

Every "doesn't do" on this list is a deliberate choice, not an oversight. The boundary is what makes the things RangerIO does do — local profiling, governed cloud handoff, edge-first compliance posture — durable. A tool that tries to be everything ends up being weaker at the parts that matter most.

If RangerIO fits inside this boundary for your work, great. If it doesn't, we'd rather you find that out from this page than three weeks into a deployment.


14. Technical FAQ

This is the technical FAQ — focused on architecture, behavior, and how the product works under the hood. For commercial / product / pricing FAQ, see the separate Product FAQ on the website.

General Architecture

Is RangerIO offline? The data engine and Privacy Scan run entirely on your laptop. The Stage 2 connector — which lets you bring your existing AI tools into the loop — is opt-in. You decide whether and when anything leaves the workstation, and Safe Handoff masks sensitive values before any cloud round-trip.

Does RangerIO send telemetry? No analytics, no telemetry, no phone-home. Crash reporting is opt-in and local-only.

Does RangerIO need an internet connection? Only for downloading the installer, downloading optional models, and (if you've enabled it) the Stage 2 connector. Day-to-day work does not require connectivity.

What's the architecture in one paragraph? A FastAPI backend on localhost:9000 runs the data engine, intelligence layer, and connector surface. The frontend is a local web app on localhost:5173 (or packaged into an Electron desktop wrapper). DuckDB handles staging and analytics queries. SQLite stores metadata. LanceDB stores vector embeddings for hybrid retrieval. fastembed generates embeddings via ONNX. GLiNER2 (also ONNX) handles entity extraction. llama-cpp-python runs the local LLM on Full SKU.

SKUs and Editions

What's the difference between Full and Lite? Full bundles a local LLM runtime so RangerIO can answer locally with no external model needed. Lite skips the bundled model and lets you bring your own — local (GGUF on disk) or cloud (via the Stage 2 connector). The data engine is identical in both.

Can I switch from Lite to Full later? Yes. Install the Full package over your existing Lite install — your projects, dossiers, and overrides are preserved.

PII Detection

What entities can RangerIO detect? A combined entity-extraction pass covers names, addresses, contact details, government IDs, financial identifiers (account numbers, credit cards), dates that re-identify, locations, organizations, and counterparties — across both structured columns and narrative text.

What's the sanity gate? A filter that drops findings without structural corroboration. For example: a column of dollar amounts may match the digit pattern of credit cards, but it lacks the column-name signals, the validator pass-rate, and the density that a real card column would show. The sanity gate drops it. Without the gate, the Privacy Scan produces unusable noise on financial data.

Are PII findings tiered? Yes — confirmed, critical, suspected, heuristic, dismissed, and user-critical. Compliance reports only count compliance-grade tiers (confirmed, critical, user_critical). Heuristic and dismissed findings do not affect compliance counts.

Do my overrides survive a rescan? Yes. The override schema persists across rescans and across model upgrades. New findings discovered on rescan show up as suspected for review.

The Dossier

What's in the dossier? Source metadata, schema and inferred types, per-column statistics, quality profile (normalization-only — format, encoding, whitespace, case, categorical drift), entity profile, PII findings, domain extensions where applicable, and a status (ready / degraded / blocked).

Why "layered"? The dossier is built in layers — source metadata, profile, domain extensions, projection views — so individual layers can rebuild independently. A PII override change rebuilds only the affected layer; a model upgrade can refresh the dossier's intelligence without re-importing the data.

How does workbook aggregation work? For Excel multi-sheet files, each sheet gets its own dossier and the workbook gets a parent dossier that aggregates findings: quality issues are summed, entities are unioned, PII findings are merged with per-sheet attribution preserved, and per-domain findings are kept separately so a workbook with both finance and HR sheets surfaces both.

Asking Questions

How does the question router work? Three handlers, fastest to most expensive: direct lookup against the dossier (counts, schema queries), Smart SQL against the local analytics engine for queries that map cleanly, and grounded advisory against the LLM for analytical or interpretive questions. RangerIO picks; you don't.

What does "grounded" mean technically? The LLM never sees the raw question alone. The grounded prompt prepends the row count, column list, dossier summary, and (when relevant) a sample of rows as authoritative facts. Explicit instructions block the model from inventing different numbers.

What happens when the LLM can't answer? RangerIO returns a clear "I don't have enough to answer that" rather than fabricating. This is the same readiness-contract pattern reports use, applied to chat.

Cloud AI and the Connector

What does the connector expose? A vendor-neutral surface that exposes dossier summaries, schema, quality profiles, the Privacy Scan summary, and project-level metadata. Raw rows are not exposed. Sensitive values are masked before any cloud round-trip.

What clients are supported? Stage 2.0 — Claude Desktop, shipping today, verified end-to-end. Stage 2.1 — ChatGPT, on the roadmap. Stage 2.2 — Perplexity, on the roadmap. They ship on the same vendor-neutral connector surface.

How is "vendor-neutral" enforced? A build check fails if any client-specific brand names leak into the connector code. This isn't a marketing posture — it's a CI gate.

What does Safe Handoff actually do? Masks sensitive values into stable stand-ins keyed by an engagement-scoped salt, sends the masked content to the cloud, and restores real values when the answer comes back. The salt rotates on engagement close or every 90 days, whichever is sooner. The stand-in-to-real-value mapping never leaves the workstation.

Hardware and Performance

How does the hardware tier system work? On first launch, RangerIO measures CPU, RAM, GPU vendor, and VRAM, then assigns one of five tiers. Batch sizes, chunk sizes, and model selection adapt automatically. The classification is frozen for the session; you can rerun it after a hardware change.

What if I don't have enough memory? A resource-readiness gate at startup offers three options: retry after closing apps, fall back to a smaller model, or skip the AI layer for now and use the data engine only. The data engine never blocks on model availability.

What models can I run? Full SKU bundles a tier-appropriate model. Any GGUF-format model on disk is loadable through the model picker. The provider abstraction layer means swapping models doesn't require code changes.

Does RangerIO use CUDA / Metal / Vulkan? On Windows with NVIDIA: CUDA. On Apple Silicon: Metal. On Windows with AMD: Vulkan. On CPU-only systems: standard CPU inference. Detection is automatic.

Data Engine

What database does RangerIO use under the hood? DuckDB for staging and analytics queries, SQLite for application metadata, LanceDB for vector embeddings.

How are documents parsed? PDFs and DOCX files are parsed with structure preserved — headings, tables, page numbers — so retrieval can attribute back to the source location.

What's the largest file RangerIO can handle? Ingestion is streamed, so file size is bounded by disk and analytics-engine performance more than by RAM. Typical limits are 10 GB+ on a Tier 3+ machine.

Live Query vs. Import Table? Live Query runs against the source database on every question — best for frequently-changing data. Import Table copies the data into RangerIO's local store — best for stable data and faster analytics.

Reports and Exports

What's a "consumer contract" on a report? Each report type declares the dossier fields it needs. RangerIO checks the dossier before rendering and returns ready, degraded, or blocked. Reports never silently substitute hallucinated content for missing data.

Can I export the dossier? Yes — Markdown export is available from any source. Sensitive values are redacted by default; you choose what's included.

Stage 2 / MCP

What protocol does the connector speak? The Model Context Protocol (MCP) — an open, vendor-neutral standard for connecting tools to AI clients. This is why the same connector surface works for Claude Desktop, ChatGPT, and Perplexity; the differences are in client-side configuration, not in the surface itself.

Does the connector run as a separate process? Yes. RangerIO ships a sidecar that handles connector traffic. The main API and the connector run independently and communicate through shared local resources.


Doc maintenance note This is the public 101-201 user guide. - Voice: professional + plain English. No regulatory claims we can't stamp. - No "100% offline" or "never leaves your machine" — the Stage 2 connector is part of the product. - For operational reference (settings catalog, advanced configuration, support flows), see the separate in-product detailed guide. - When adding a capability: update the relevant page AND the FAQ. When deprecating: remove from this guide; do not leave stale references.