Home / Tools / Curated datasets
find_data_sources
Find Data Sources
What it answers
Search the Vlada substrate catalog by keyword. Returns the matching data sources (commercial TiC, hospital HPT, Medicare, drug pricing, provider identity, curated marts, provenance registries) ranked by relevance. Compact by default (identity + queryable + claim ceiling + the exact describe_data_source call per hit); pass verbosity='full' for the previous full source_product_card per result. Paginates with limit (default 25, max 200) + offset; the response carries total_count, returned_count, and offset. Optional filters: {domain, stage, status, limit}. Agents call this first to discover what data Vlada has. Since 2026-09-02 the answer leads with DATA POINTS (`data_points`: the measure/dimension asked for, with its source, door, table.column and must-read trap) and the table cards follow as physical locations. Every hit carries `source_code` -- the one public code, which describe_data_source accepts verbatim -- alongside the table's `source_fqn`. Points and tables are ranked by ONE ranker (the authored data-point scorer). Retired tables are never returned; they appear only as `supersedes` on the live successor's card. BOUNDED (2026-09-18): the default call is compact and paged -- both the table cards and the data points come back a page at a time, each with its own total and next offset -- so a default call can never truncate. Pass verbosity='full' for the whole point record and the full source_product_card per hit.
Inputs
filterslimitmax_pointsoffsetqueryverbosityworldCall it
From an agent: connect the Vlada MCP once (one click for Claude, ChatGPT, Cursor, VS Code) and ask in plain English; the agent selects find_data_sources when the question fits. From code: the same tool over REST with an API key. The schema endpoint needs no key.
curl -X POST https://api.vladahealth.com/v1/tools/find_data_sources \
-H "Authorization: Bearer $VLADA_API_KEY" \
-H "Content-Type: application/json" \
-d '{}'curl https://api.vladahealth.com/v1/tools/find_data_sources/schema # the JSON schema, no auth # MCP endpoint (Streamable HTTP): https://mcp.vladahealth.com/mcp
What comes back
Typed rows plus provenance on every answer: the public source file, its vintage, the methodology, and a response_hash you can replay. Prove and replay tools turn any number into a re-runnable receipt. A number the data cannot support comes back as “not in the data”, never as zero.
Output schema
{
"$defs": {
"Quality": {
"description": "Row-level quality signal. All flags precomputed at build time\n(stored on the gold mart) and passed through here.",
"properties": {
"is_outlier": {
"default": false,
"title": "Is Outlier",
"type": "boolean"
},
"is_ghost_candidate": {
"default": false,
"title": "Is Ghost Candidate",
"type": "boolean"
},
"n_similar_rates": {
"default": 0,
"title": "N Similar Rates",
"type": "integer"
},
"confidence": {
"default": 1,
"title": "Confidence",
"type": "number"
}
},
"title": "Quality",
"type": "object"
},
"Source": {
"description": "Cell-level provenance. Every served value traces back here.\n\n`snapshot_id` is the Iceberg snapshot the tool read from. Two calls\nat the same snapshot MUST produce byte-identical responses —\nthat's the content-addressed caching guarantee.",
"properties": {
"table": {
"title": "Table",
"type": "string"
},
"rate_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Rate Id"
},
"source_files": {
"items": {
"type": "string"
},
"title": "Source Files",
"type": "array"
},
"snapshot_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Snapshot Id"
}
},
"required": [
"table"
],
"title": "Source",
"type": "object"
}
},
"properties": {
"value": {
"default": null,
"title": "Value"
},
"unit": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Unit"
},
"vintage": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Vintage"
},
"source": {
"$ref": "#/$defs/Source"
},
"methodology": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Methodology"
},
"quality": {
"$ref": "#/$defs/Quality"
},
"invariants_applied": {
"items": {
"type": "string"
},
"title": "Invariants Applied",
"type": "array"
},
"caveats": {
"items": {
"type": "string"
},
"title": "Caveats",
"type": "array"
},
"response_hash": {
"default": "",
"title": "Response Hash",
"type": "string"
},
"semantic_version": {
"default": 1,
"title": "Semantic Version",
"type": "integer"
},
"tool_version": {
"default": "unknown",
"title": "Tool Version",
"type": "string"
},
"explanation": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Explanation"
},
"status": {
"default": "complete",
"title": "Status",
"type": "string"
},
"refusal": {
"anyOf": [
{
"additionalProperties": true,
"type": "object"
},
{
"type": "null"
}
],
"default": null,
"title": "Refusal"
},
"failure": {
"anyOf": [
{
"additionalProperties": true,
"type": "object"
},
{
"type": "null"
}
],
"default": null,
"title": "Failure"
},
"as_of": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "As Of"
},
"freshness_state": {
"default": "unknown",
"title": "Freshness State",
"type": "string"
}
},
"title": "CatalogResponse",
"type": "object"
}Sources behind it
Known limits
No gaps recorded for this tool. Absence of a recorded gap is not a claim of complete coverage; the answer itself says what it covers.
As of the 2026-09-19 build of the served surface · machine-readable catalog · the live server may run a different version; the schema endpoint above is authoritative for what is deployed.