openalex-mcp-server

v0.8.0 pre-1.0

Search the OpenAlex catalog — 270M+ works, 90M+ authors, 100K+ sources.

openalex.caseyjhand.com/mcp
claude mcp add --transport http openalex-mcp-server https://openalex.caseyjhand.com/mcp
codex mcp add openalex-mcp-server --url https://openalex.caseyjhand.com/mcp
{
  "mcpServers": {
    "openalex-mcp-server": {
      "url": "https://openalex.caseyjhand.com/mcp"
    }
  }
}
gemini mcp add --transport http openalex-mcp-server https://openalex.caseyjhand.com/mcp
{
  "mcpServers": {
    "openalex-mcp-server": {
      "command": "bunx",
      "args": [
        "mcp-remote",
        "https://openalex.caseyjhand.com/mcp"
      ],
      "env": {
        "OPENALEX_API_KEY": "your-openalex-api-key"
      }
    }
  }
}
{
  "mcpServers": {
    "openalex-mcp-server": {
      "type": "http",
      "url": "https://openalex.caseyjhand.com/mcp"
    }
  }
}
curl -X POST https://openalex.caseyjhand.com/mcp \
  -H "Content-Type: application/json" \
  -H "Accept: application/json, text/event-stream" \
  -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-11-25","capabilities":{},"clientInfo":{"name":"curl","version":"1.0.0"}}}'

Tools

5

openalex_resolve_name

open-world

Resolve a name or an identifier to an OpenAlex ID. ALWAYS use this before filtering by entity — names are ambiguous, IDs are not. A name returns up to 10 autocomplete matches with disambiguation hints. An identifier — OpenAlex ID, DOI, ORCID, ROR, PMID, or ISSN, bare or in URL form — resolves directly to the one record it addresses, and needs no entity_type. A PMCID is recognized as well, bare or as a PubMed Central URL, but OpenAlex indexes no PMCIDs, so it resolves nothing — pass the work's PMID or DOI instead.

read
invocation
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "openalex_resolve_name",
    "arguments": {
      "query": "<query>"
    }
  }
}
schema
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "type": "object",
  "properties": {
    "entity_type": {
      "description": "Entity type to search. Omit for cross-entity search (useful when entity type is unknown). Not applied when `query` is an identifier — an identifier determines its own entity type.",
      "type": "string",
      "enum": [
        "works",
        "authors",
        "sources",
        "institutions",
        "topics",
        "keywords",
        "publishers",
        "funders"
      ]
    },
    "query": {
      "type": "string",
      "minLength": 1,
      "description": "Name or partial name to resolve. Also accepts an identifier, bare or in URL form — OpenAlex ID (\"W2741809807\", \"F4320332161\"), DOI (\"10.1038/nature12373\"), ORCID (\"0000-0002-1825-0097\"), ROR (\"https://ror.org/00hx57361\"), PMID (\"12345678\" or \"https://pubmed.ncbi.nlm.nih.gov/12345678\"), ISSN (\"1234-5678\") — which resolves straight to that one record instead of running a name search. A keyword URL (\"https://openalex.org/keywords/groundwater\") resolves the same way; a bare keyword slug reads as a name and runs a name search, which finds it too. A PMCID (\"PMC1234567\" or a PubMed Central URL) is recognized but OpenAlex indexes no PMCIDs, so it resolves nothing — pass the work's PMID or DOI instead."
    },
    "filters": {
      "description": "Narrow autocomplete results with filters. Example: restrict to a specific country or publication year range. Applies to name queries only — an identifier already addresses a single record.",
      "type": "object",
      "propertyNames": {
        "type": "string"
      },
      "additionalProperties": {
        "type": "string"
      }
    }
  },
  "required": [
    "query"
  ],
  "additionalProperties": false
}
view source ↗

openalex_search_entities

open-world

Search, filter, sort, or retrieve by ID. Covers all OpenAlex entity types (works, authors, sources, institutions, topics, keywords, publishers, funders). Pass `id` to retrieve a single entity. Otherwise, use `query` and/or `filters` for discovery. Supports keyword search with boolean operators, exact phrase matching, and AI semantic search. Use openalex_resolve_name to resolve names to IDs before filtering. Searches and ID lookups return a curated set of fields by default; pass `select` to override with specific fields, or `["*"]` for the full record. Responses cap at 64,000 bytes per surface unless the least a call can return is larger (`over_budget`); `omitted` and `windows` give the calls that continue a cut, and `slice` pages a long array.

read
invocation
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "openalex_search_entities",
    "arguments": {
      "entity_type": "<entity_type>"
    }
  }
}
schema
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "type": "object",
  "properties": {
    "entity_type": {
      "type": "string",
      "enum": [
        "works",
        "authors",
        "sources",
        "institutions",
        "topics",
        "keywords",
        "publishers",
        "funders"
      ],
      "description": "Type of scholarly entity to search."
    },
    "id": {
      "description": "Retrieve a single entity by ID. Supports: OpenAlex ID (\"W2741809807\"), DOI (\"10.1038/nature12373\"), ORCID (\"0000-0002-1825-0097\"), ROR (\"https://ror.org/00hx57361\"), PMID (\"12345678\" or \"https://pubmed.ncbi.nlm.nih.gov/12345678\"), ISSN (\"1234-5678\"). Keywords are identified by slug rather than a native ID — pass either the slug (\"groundwater\") or the URL a search returns (\"https://openalex.org/keywords/groundwater\"). A PMCID is recognized too, bare (\"PMC1234567\") or as a PubMed Central URL, but OpenAlex indexes no PMCIDs, so it resolves nothing — pass the work's PMID or DOI instead. When provided, `query`, `search_mode`, `filters`, `sort`, `sample`, and `seed` are not applied — the returned record is the entity at that ID regardless of them, and the response `notice` names any you passed. `select` still applies: the curated per-entity-type default is returned unless you pass `select` (use `[\"*\"]` for the complete record). An array too long for the response budget comes back as a window (`windows`); page the rest with `slice`. To filter, drop `id` and search. Use openalex_resolve_name to find the ID if unknown.",
      "type": "string",
      "minLength": 1
    },
    "query": {
      "description": "Text search query. Supports boolean operators (AND, OR, NOT), quoted phrases (\"exact match\"), wildcards (machin*), fuzzy matching (machin~1), and proximity (\"climate change\"~5). Omit for filter-only queries — an empty string is rejected, since a blank search is a mistake rather than a request for the whole catalog.",
      "type": "string",
      "minLength": 1
    },
    "search_mode": {
      "default": "keyword",
      "description": "Search strategy. \"keyword\": stemmed full-text (default). \"exact\": no stemming, matches individual words (use quoted phrases for multi-word exact match). \"semantic\": AI embedding similarity over a query-dependent candidate set whose size `meta.count` reports, at ~1 req/sec, up to 50 per page, and paginated with `page` rather than `cursor`.",
      "type": "string",
      "enum": [
        "keyword",
        "exact",
        "semantic"
      ]
    },
    "filters": {
      "description": "Filter criteria as field:value pairs. AND across fields (multiple keys). OR within field: pipe-separate (\"us|gb\"). NOT: prefix \"!\" (\"!us\"). Range: \"2020-2024\". Comparison: \">100\", \"<50\". AND within same field: \"+\"-separate. Two keys that resolve to the same upstream field (an alias and its canonical name, e.g. `year` and `publication_year`) are both applied and AND'd, so they narrow rather than override each other. Use OpenAlex IDs (not names) for entity filters — resolve names first. Common keys: `openalex` (filter by entity ID, e.g. {\"openalex\": \"W123|W456\"}), `cites` (works citing a given work), `publication_year` (range \"2020-2024\"), `authorships.author.id`, `type`, `is_oa`.",
      "type": "object",
      "propertyNames": {
        "type": "string"
      },
      "additionalProperties": {
        "type": "string"
      }
    },
    "sort": {
      "description": "Sort field. Prefix with \"-\" for descending. Comma-separate for a multi-key sort, applied left to right, with the \"-\" prefix set per key (\"-publication_year,cited_by_count\" sorts by year descending, then citations ascending). Common: \"cited_by_count\", \"-publication_date\", \"-relevance_score\" (default when query present). Note: when combined with a keyword query, an explicit sort overrides relevance ranking entirely — top results may be highly cited but only tangentially on-topic. Use \"-relevance_score\" or omit sort to keep the most relevant results first. \"-relevance_score\" requires an active search via \"query\" or a \"filter:search\" filter — passing it without one will fail. Not combinable with `sample` — a search passing both is rejected.",
      "type": "string"
    },
    "select": {
      "description": "OpenAlex top-level field names to return. Always returned: `id`, `display_name` — additional fields you list are appended. A curated default per entity type applies to both searches and single-entity (`id`) lookups; pass field names to override it, or `[\"*\"]` to retrieve the complete record (every field). Only top-level fields project, so a nested value is requested by its parent object: bibliometrics (`h_index`, `i10_index`, `2yr_mean_citedness`) live under `summary_stats` on authors, sources, institutions, publishers, and funders, and naming a leaf returns that object. Invalid field names produce an error identifying the rejected field. Large projections may be cut to fit the response budget (`omitted`, `windows`). Not applied with `slice`. Example: [\"doi\", \"authorships\", \"primary_topic\"].",
      "type": "array",
      "items": {
        "type": "string"
      }
    },
    "per_page": {
      "default": 25,
      "description": "Results per page (1-100). Default 25. Semantic search caps at 50 — when search_mode=\"semantic\", set per_page ≤ 50 (also subject to a 1 req/sec rate limit upstream). The cap applies to searches only; an `id` lookup returns its one record regardless of both. A budget-cut page returns fewer (`omitted`).",
      "type": "integer",
      "minimum": 1,
      "maximum": 100
    },
    "cursor": {
      "description": "Pagination cursor from a previous response. Pass to get the next page. Omit it on the first call — an empty string is rejected, since a supplied-but-blank cursor is a caller mistake rather than a request for page 1. Keyword and exact modes only — semantic search walks its candidates with `page`, and a `cursor` sent with it is rejected.",
      "type": "string",
      "minLength": 1
    },
    "page": {
      "description": "Page number (1-based) for semantic search, the one mode that paginates with `page` instead of `cursor`. Semantic search ranks a query-dependent candidate set whose size `meta.count` reports, so the last reachable page is ceil(meta.count / per_page) — e.g. page 14 for a count of 70 with per_page=5. Passing it under any other search_mode is rejected.",
      "type": "integer",
      "minimum": 1,
      "maximum": 9007199254740991
    },
    "sample": {
      "description": "Return a random sample of this many entities matching the filters (1-100). Single page only — neither `cursor` nor `page` pagination applies to sampling, and a search that passes either alongside it is rejected. Keyword and exact modes only: OpenAlex does not sample a semantic search, so `sample` with search_mode \"semantic\" is rejected. Cannot be combined with `sort` — a sample has no order, and a search passing both is rejected. Overrides `per_page`. Useful for unbiased exploration: spot-checking filter correctness, stratified review prompts, or generating exploration sets without bias toward most-cited.",
      "type": "integer",
      "minimum": 1,
      "maximum": 100
    },
    "seed": {
      "description": "Deterministic seed for `sample`. Same seed + same filters = same results — pass when reproducibility matters. Has no effect without `sample`, and a search that passes it alone is rejected.",
      "type": "string"
    },
    "slice": {
      "description": "Page one array of the record at `id`: returns `id`, `display_name`, and `field` from `offset` onward, as many elements as fit the response budget (at least one), with the window and its `next` call in `windows`. Requires `id`; replaces `select`. An offset at or past the end returns an empty window. Walks any partial array, including a possibly capped 100-authorship list.",
      "type": "object",
      "properties": {
        "field": {
          "type": "string",
          "description": "Top-level array field to page, such as \"authorships\" or \"referenced_works\". Leave it empty to skip slicing."
        },
        "offset": {
          "default": 0,
          "description": "0-based index of the first element to return.",
          "type": "integer",
          "minimum": 0,
          "maximum": 9007199254740991
        }
      },
      "required": [
        "field",
        "offset"
      ],
      "additionalProperties": false
    }
  },
  "required": [
    "entity_type",
    "search_mode",
    "per_page"
  ],
  "additionalProperties": false
}
view source ↗

openalex_get_citation_graph

open-world

Walk the citation graph one hop from a seed work. Direction picks the edge: incoming citations (`cites`), the seed's own references (`cited_by`), or OpenAlex's algorithmically-related works (`related_to`). Note: `direction` follows OpenAlex's filter convention, which inverts the common English reading — `cites` returns works that cite the seed; `cited_by` returns works the seed cites. Results use the works schema; combine with filters/sort to narrow further. Responses cap at 64,000 bytes per surface unless the least a call can return is larger (`over_budget`); `omitted` and `windows` give the openalex_search_entities calls that continue a cut.

read
invocation
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "openalex_get_citation_graph",
    "arguments": {
      "seed_id": "<seed_id>",
      "direction": "<direction>"
    }
  }
}
schema
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "type": "object",
  "properties": {
    "seed_id": {
      "type": "string",
      "minLength": 1,
      "description": "Seed work identifier. Accepts OpenAlex ID (\"W2741809807\"), DOI (\"10.1038/nature12373\" or full URL), or PMID (\"12345678\" or \"https://pubmed.ncbi.nlm.nih.gov/12345678\"). A PMCID is recognized too, bare or as a PubMed Central URL, but OpenAlex indexes no PMCIDs, so it resolves nothing — pass the work's PMID or DOI instead. Use openalex_resolve_name first if you only have a title."
    },
    "direction": {
      "type": "string",
      "enum": [
        "cites",
        "cited_by",
        "related_to"
      ],
      "description": "\"cites\": works that cite seed_id (incoming citations). \"cited_by\": works that seed_id cites (its reference list). \"related_to\": OpenAlex algorithmically-related works (~8-30 typical, may be empty for less-cited seeds)."
    },
    "filters": {
      "description": "Additional filters to narrow the graph, same syntax as openalex_search_entities. Example: publication_year=\">2020\", is_oa=\"true\". Do not include cites/cited_by/related_to, nor an alias of one such as cited_works — those keys are set by the `direction` parameter.",
      "type": "object",
      "propertyNames": {
        "type": "string"
      },
      "additionalProperties": {
        "type": "string"
      }
    },
    "sort": {
      "description": "Sort field. Prefix with \"-\" for descending. Comma-separate for a multi-key sort, applied left to right, with the \"-\" prefix set per key (\"-publication_year,cited_by_count\" sorts by year descending, then citations ascending). Common: \"cited_by_count\", \"-publication_date\". Default is OpenAlex relevance.",
      "type": "string"
    },
    "select": {
      "description": "OpenAlex work field names to return. Always returned: id, display_name. Defaults to the curated works select if omitted. Large projections may be cut to fit the response budget (`omitted`, `windows`).",
      "type": "array",
      "items": {
        "type": "string"
      }
    },
    "per_page": {
      "default": 25,
      "description": "Results per page (1-100). Default 25. A budget-cut page returns fewer (`omitted`).",
      "type": "integer",
      "minimum": 1,
      "maximum": 100
    },
    "cursor": {
      "description": "Pagination cursor from a previous response. Pass to get the next page. Omit it on the first call — an empty string is rejected, since a supplied-but-blank cursor is a caller mistake rather than a request for the first page.",
      "type": "string",
      "minLength": 1
    }
  },
  "required": [
    "seed_id",
    "direction",
    "per_page"
  ],
  "additionalProperties": false
}
view source ↗

openalex_describe_fields

List valid field names for an OpenAlex entity type and context (filter, group_by, or select). Use proactively before constructing a filter or group_by to avoid invalid-field 400 errors. Pass `query` to rank the list by name similarity — useful when you have a partial or guessed field name. Ranking never drops a field: the full list comes back either way.

read
invocation
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "openalex_describe_fields",
    "arguments": {
      "entity_type": "<entity_type>",
      "context": "<context>"
    }
  }
}
schema
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "type": "object",
  "properties": {
    "entity_type": {
      "type": "string",
      "enum": [
        "works",
        "authors",
        "sources",
        "institutions",
        "topics",
        "keywords",
        "publishers",
        "funders"
      ],
      "description": "OpenAlex entity type to list fields for."
    },
    "context": {
      "type": "string",
      "enum": [
        "filter",
        "group_by",
        "select"
      ],
      "description": "Field usage context. \"filter\": fields accepted in the filter param. \"group_by\": fields accepted in group_by — a subset of the filter set that leaves out what OpenAlex refuses to aggregate (raw dates, *.search operators, decimal scores, display_name, and external-ID fields among them). \"select\": fields accepted in select."
    },
    "query": {
      "description": "Optional partial or guessed field name to sort results by similarity. Pass the field you tried (e.g. \"funder\") to get the closest matches first. The complete field list is returned either way — a query reorders it, it does not filter it, so a nested value's parent object (e.g. `summary_stats` for \"h_index\") is still reachable further down.",
      "type": "string"
    }
  },
  "required": [
    "entity_type",
    "context"
  ],
  "additionalProperties": false
}
view source ↗

Prompts

2

Guides a systematic literature search: formulate query, search, filter, analyze citation network, synthesize findings.

  • topicrequired — Research topic or question to review (e.g., "CRISPR off-target effects in human cell lines").
  • scope — "narrow": focused on specific question. "broad": survey of the field. Defaults to "narrow" if omitted.

Analyzes the research landscape for a topic: volume trends, top authors/institutions, open access rates, funding sources.

  • topicrequired — Research area to analyze (e.g., "single-cell RNA sequencing").