Claude skill 06 · openfda-search

openFDA search

Search FDA's public device data through the openFDA API — adverse event reports (MAUDE), recalls and enforcement reports, 510(k) clearances, PMA approvals and supplements, device classification, registration and listing, and the UDI database — for a device, a manufacturer, a product code, or a failure mode; pull every matching record into files, not screenshots; count by year, by problem, by model; and write a search record that states the query, the date, and the limits of the data.

SKILL.md

The instructions, as Claude reads them.

One folder, one file. Download the .skill file and add it in Claude’s skills settings, or unzip it into your skills folder. Our guide 06 explains the method and its limits.

The skill states what the records say and where. It does not form opinions on cause or liability. Check every citation before a document leaves your team.

FDA publishes its device databases through one API, api.fda.gov. The web search pages show one record at a time; the API returns every matching record as data, so the search can be repeated, counted, and cited. This skill pulls the data into files, writes the counts, and records the query, because a search that cannot be repeated by the other side is not evidence of anything.

The data has known limits, and every output must say so. MAUDE is a passive, largely voluntary reporting system: a count of reports is not a count of events, one event can appear several times, reports can be wrong or incomplete, and the absence of a report is not the absence of a failure. The output states what the records say and how many there are in the database as of the day of the search. It draws no conclusion about rates or about cause.

1. Inputs

  • The device: brand name, manufacturer, model or catalog number, product code if known, and the failure mode or component of interest. Spell names the way FDA spells them: search the UDI and classification endpoints first to learn the product code and the manufacturer's names as registered (a manufacturer often appears under several names).
  • The period: default from January 1 of the year the device was first cleared to today; narrow it only when the user asks.
  • An API key when available (OPENFDA_API_KEY in the environment; the user gets one free at open.fda.gov/apis/authentication). Without a key the limit is 240 requests per minute and 1,000 per day per IP; with a key, 240 per minute and 120,000 per day. Never ask the user to paste the key into the chat; read it from the environment.

2. The endpoints

All under https://api.fda.gov/device/. Each takes search=, count=, limit= (max 1,000), skip= (max 25,000), sort=, and api_key=.

Endpoint What it holds Fields that matter most
event.json Adverse event reports (MAUDE), 1991 to date, refreshed weekly report_number, mdr_report_key, date_received, date_of_event, event_type (Death / Injury / Malfunction / Other), device.brand_name, device.generic_name, device.manufacturer_d_name, device.model_number, device.catalog_number, device.lot_number, device.device_report_product_code, product_problems, patient.sequence_number_outcome, mdr_text.text (narratives: event description and manufacturer narrative), source_type, remedial_action, device.device_operator, openfda.* (harmonized product code, device name, regulation)
recall.json Device recalls (medical device recall database) res_event_number, product_res_number, recalling_firm, product_description, product_code, root_cause_description, event_date_initiated, event_date_posted, recall_status, firm_fda_model_number, k_numbers, pma_numbers, action, reason_for_recall
enforcement.json Enforcement reports (classified recalls, I / II / III) recall_number, classification, status, recalling_firm, product_description, reason_for_recall, recall_initiation_date, report_date, code_info, distribution_pattern
510k.json Premarket notifications k_number, applicant, device_name, product_code, decision_date, decision_description, clearance_type, statement_or_summary, third_party_flag
pma.json Premarket approvals and supplements pma_number, supplement_number, supplement_type, supplement_reason, applicant, trade_name, generic_name, product_code, decision_date, decision_code, ao_statement
classification.json Product classification product_code, device_name, device_class, regulation_number, medical_specialty_description, submission_type_id, definition
registrationlisting.json Establishment registration and device listing registration.name, registration.fei_number, products.product_code, products.openfda.device_name, proprietary_name
udi.json Global UDI database (GUDID) brand_name, company_name, version_or_model_number, catalog_number, identifiers.id, product_codes.code, gmdn_terms.name, device_description, mri_safety, is_single_use

Query syntax that is easy to get wrong: - Phrases in quotes: search=device.brand_name:"infusion pump". Operators AND, OR in capitals; + is the space in a URL. - Exact matching for counts and for strings with punctuation: count=device.manufacturer_d_name.exact. - Date ranges: date_received:[20190101+TO+20241231] (dates are YYYYMMDD). - Missing fields: _missing_:date_of_event; existence: _exists_:device.lot_number. - Narrative text: mdr_text.text:"display froze"; the text is tokenized, so a phrase search matches the words in order, not the exact string. - skip stops at 25,000. For larger sets, split the period into windows (by month or by year) until every window returns fewer than 25,000, or use sort=date_received:asc with the search_after parameter the API offers for deep paging. Never sample; pull the whole set.

3. Search in the right order

  1. Identify the device in FDA's own terms. classification.json by name to get product_code, device_class, regulation_number. udi.json and 510k.json / pma.json to get the names FDA has for the manufacturer and the model, every spelling (count=device.manufacturer_d_name.exact on event.json for the product code shows the variants). Write them down; every later query uses them.
  2. The regulatory record: every 510(k) and PMA (with supplements) for the manufacturer and product code; decision dates and types. This is the stream A material for the device-timeline skill.
  3. Recalls and enforcement: by recalling_firm variants and by product_code; then read product_description and firm_fda_model_number to keep only the model at issue and its sister models.
  4. Adverse events: by product code and manufacturer variants, over the period; then narrow by brand name, model, and narrative terms for the failure mode. Pull the full records, including every mdr_text entry.
  5. Counts before and after each narrowing step, so the summary can say how many reports exist for the product code, for the manufacturer, for the model, and for the failure mode terms, with the query for each.

4. Pull the data

Use a script; never read result pages in the browser one by one. The script below pulls every record for a query into one JSON-lines file, with the date windows and the request log beside it. Run it with python3 -I.

import os, sys, json, time, csv, datetime as dt, urllib.parse, urllib.request

BASE = 'https://api.fda.gov/device/'
KEY = os.environ.get('OPENFDA_API_KEY', '')

def get(endpoint, params, log):
    q = dict(params)
    if KEY: q['api_key'] = KEY
    url = BASE + endpoint + '?' + urllib.parse.urlencode(q, safe='+:[]"')
    for attempt in range(5):
        try:
            with urllib.request.urlopen(url, timeout=60) as r:
                data = json.load(r)
            log.append((dt.datetime.utcnow().isoformat(), url.replace(KEY, 'KEY') if KEY else url, data.get('meta', {}).get('results', {}).get('total')))
            return data
        except urllib.error.HTTPError as e:
            if e.code == 404:          # no matches
                log.append((dt.datetime.utcnow().isoformat(), url.replace(KEY, 'KEY') if KEY else url, 0))
                return {'results': [], 'meta': {'results': {'total': 0}}}
            if e.code == 429:          # rate limit
                time.sleep(15); continue
            raise
    raise RuntimeError('gave up: ' + url)

def pull(endpoint, search, date_field, start, end, out, log):
    """Every record for `search` between start and end (YYYYMMDD), split by windows so skip stays under 25,000."""
    total = 0
    windows = [(start, end)]
    with open(out, 'a') as fh:
        while windows:
            a, b = windows.pop()
            s = f'{search}+AND+{date_field}:[{a}+TO+{b}]' if search else f'{date_field}:[{a}+TO+{b}]'
            meta = get(endpoint, {'search': s, 'limit': 1}, log)['meta']['results']['total']
            if meta > 25000 and a != b:                       # split the window
                da, db = dt.datetime.strptime(a, '%Y%m%d'), dt.datetime.strptime(b, '%Y%m%d')
                mid = da + (db - da) / 2
                windows += [(a, mid.strftime('%Y%m%d')), ((mid + dt.timedelta(days=1)).strftime('%Y%m%d'), b)]
                continue
            for skip in range(0, meta, 1000):
                data = get(endpoint, {'search': s, 'limit': 1000, 'skip': skip, 'sort': f'{date_field}:asc'}, log)
                for rec in data['results']:
                    fh.write(json.dumps(rec) + '\n'); total += 1
                time.sleep(0.3)
    return total

if __name__ == '__main__':
    endpoint, search, date_field, start, end, out = sys.argv[1:7]
    log = []
    n = pull(endpoint, search, date_field, start, end, out, log)
    with open(out + '.log.csv', 'a', newline='') as fh:
        csv.writer(fh).writerows(log)
    print(n, 'records ->', out)

Example: python3 -I pull.py event.json 'device.device_report_product_code:FRN+AND+device.manufacturer_d_name:"ACME"' date_received 20150101 20261231 acme-frn-events.jsonl

Then flatten the records you need into CSV with the columns the question needs (one row per report; for MAUDE, one row per mdr_text entry when the narratives are the point), and keep the JSONL as the record of what the API returned.

5. Count and read

  • Counts: by year (count=date_received per year, or compute from the pulled set), by event_type, by product_problems.exact, by device.model_number.exact, by patient.sequence_number_outcome.exact. Give each count with its query and the date of the search.
  • Duplicates: group by mdr_report_key and by (date_of_event, device.lot_number, narrative start) to find the same event reported by the manufacturer and by the user facility; report both counts, the raw and the deduplicated, and say how you grouped.
  • Narratives: read every narrative for the failure mode set; code them with a short, defined code list (as the complaint-reader skill does); quote short phrases verbatim with the report_number.
  • Cross-reference: recalls to 510(k) and PMA numbers (k_numbers, pma_numbers); events to recalls by model and period; the device at issue to its sister models by product code and firm.

6. Write the output

  1. openfda-search-record.md — the method record: the date and time of the search; the endpoints; every query string as sent (key removed); the totals returned; the date windows used; the API's own meta.last_updated for each endpoint; the limits of the data (the MAUDE caveats in FDA's own words from the database page, cited). This file is what makes the search repeatable, and it goes with the output wherever it goes.
  2. The data: *.jsonl as returned; *.csv flattened; counts.csv with query, count, date.
  3. openfda-summary.md: the device as FDA identifies it (product code, class, regulation, clearances and approvals with dates and numbers); recalls and enforcement reports for the model and its sisters, with dates, classes, and FDA's stated reason; adverse event counts by year and type for the product code, the manufacturer, the model, and the failure mode terms, raw and deduplicated; the narratives for the failure mode, coded, with report numbers; what the search could not resolve (manufacturer name variants not checked, narratives too short to code, product codes that changed over time). Every number carries its query reference from the search record.

Citation form for FDA records: [MAUDE 1234567-2023-00012] (the report_number), [Recall Z-1234-2023] or [Recall event 91234], [K123456], [P120034/S005], [GUDID 00812345678901].

7. Rules

  • Pull the whole set, in files; never sample, never work from the browser page.
  • FDA's names and codes first; learn the variants before counting.
  • Every count with its query, its date, and the raw and deduplicated figures.
  • The caveats go in every summary, in the first section, not in a footnote: reports are not events, counts are not rates, absence is not evidence, and the data is as of the search date.
  • No rates per device or per year of use unless the user supplies the denominator from a documented source and asks for it; then show the arithmetic.
  • No conclusion on cause, on reportability, or on what the manufacturer should have done. The summary says what FDA's databases contain. The device-timeline and complaint-reader skills take it from there.

Get in touch

Tell us about the device.

Share a brief overview of the device, the question you need answered, and any deadlines. We’ll explain how we can help and recommend the next steps.

Start a conversation