Home / Tools / Commercial payer rates
query_substrate
Query Substrate
What it answers
Run a read-only SQL SELECT directly over the Vlada cost-data substrate (Athena/Iceberg) and get the rows back, with provenance. Use this when the pre-built tools don't answer your exact question. WORKFLOW: first call find_data_sources / describe_data_source to find the table + read its card (grain, columns, and the CAVEATS you must apply — e.g. filter ghost rates, never blend across modifiers/methodologies), THEN write SQL against the table's fully-qualified name (e.g. vlada_curated.fact_hospital_chargemaster). RULES: read-only SELECT only (no INSERT/UPDATE/DDL — rejected); you may only query tables in the queryable set your account is entitled to (others are rejected); M-2 versioned tables must be filtered on is_current (injected for a single-table SELECT, otherwise rejected with the column named); results are LIMIT-bounded (max 10000), the delivered payload is bounded to ~47K chars, and the scan is cost-capped. Any bound that drops rows is reported as truncated:true with rows_returned + rows_dropped/total_rows_estimate and a caveat saying how to get the rest (aggregate in SQL, narrower predicate, ORDER BY + OFFSET paging, or submit_query for a scan that needs more time). Every result carries provenance (the engine, query id, bytes scanned, tables read). Apply the source card's methodology yourself — raw SQL does not enforce it.
Inputs
sqlmax_rowsCall it
From an agent: connect the Vlada MCP once (one click for Claude, ChatGPT, Cursor, VS Code) and ask in plain English; the agent selects query_substrate when the question fits. From code: the same tool over REST with an API key. The schema endpoint needs no key.
curl -X POST https://api.vladahealth.com/v1/tools/query_substrate \
-H "Authorization: Bearer $VLADA_API_KEY" \
-H "Content-Type: application/json" \
-d '{"sql":"<sql>"}'curl https://api.vladahealth.com/v1/tools/query_substrate/schema # the JSON schema, no auth # MCP endpoint (Streamable HTTP): https://mcp.vladahealth.com/mcp
What comes back
Typed rows plus provenance on every answer: the public source file, its vintage, the methodology, and a response_hash you can replay. Prove and replay tools turn any number into a re-runnable receipt. A number the data cannot support comes back as “not in the data”, never as zero.
Output schema
{
"$defs": {
"Quality": {
"description": "Row-level quality signal. All flags precomputed at build time\n(stored on the gold mart) and passed through here.",
"properties": {
"is_outlier": {
"default": false,
"title": "Is Outlier",
"type": "boolean"
},
"is_ghost_candidate": {
"default": false,
"title": "Is Ghost Candidate",
"type": "boolean"
},
"n_similar_rates": {
"default": 0,
"title": "N Similar Rates",
"type": "integer"
},
"confidence": {
"default": 1,
"title": "Confidence",
"type": "number"
}
},
"title": "Quality",
"type": "object"
},
"Source": {
"description": "Cell-level provenance. Every served value traces back here.\n\n`snapshot_id` is the Iceberg snapshot the tool read from. Two calls\nat the same snapshot MUST produce byte-identical responses —\nthat's the content-addressed caching guarantee.",
"properties": {
"table": {
"title": "Table",
"type": "string"
},
"rate_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Rate Id"
},
"source_files": {
"items": {
"type": "string"
},
"title": "Source Files",
"type": "array"
},
"snapshot_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Snapshot Id"
}
},
"required": [
"table"
],
"title": "Source",
"type": "object"
}
},
"properties": {
"value": {
"default": null,
"title": "Value"
},
"unit": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Unit"
},
"vintage": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Vintage"
},
"source": {
"$ref": "#/$defs/Source"
},
"methodology": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Methodology"
},
"quality": {
"$ref": "#/$defs/Quality"
},
"invariants_applied": {
"items": {
"type": "string"
},
"title": "Invariants Applied",
"type": "array"
},
"caveats": {
"items": {
"type": "string"
},
"title": "Caveats",
"type": "array"
},
"response_hash": {
"default": "",
"title": "Response Hash",
"type": "string"
},
"semantic_version": {
"default": 1,
"title": "Semantic Version",
"type": "integer"
},
"tool_version": {
"default": "unknown",
"title": "Tool Version",
"type": "string"
},
"explanation": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Explanation"
},
"status": {
"default": "complete",
"title": "Status",
"type": "string"
},
"refusal": {
"anyOf": [
{
"additionalProperties": true,
"type": "object"
},
{
"type": "null"
}
],
"default": null,
"title": "Refusal"
},
"failure": {
"anyOf": [
{
"additionalProperties": true,
"type": "object"
},
{
"type": "null"
}
],
"default": null,
"title": "Failure"
},
"as_of": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "As Of"
},
"freshness_state": {
"default": "unknown",
"title": "Freshness State",
"type": "string"
}
},
"title": "QuerySubstrateResponse",
"type": "object"
}Sources behind it
Known limits
- Raw SQL does NOT enforce a source's methodology (ghost-rate filter, modifier/methodology non-blending) — the agent must apply the caveats from the source card. For warranted, methodology-enforced numbers use the source's specialized tool.
As of the 2026-09-19 build of the served surface · machine-readable catalog · the live server may run a different version; the schema endpoint above is authoritative for what is deployed.