GuideStar / ProPublica Nonprofit Scraper Services
Nenodata's GuideStar / ProPublica Nonprofit Scraper service turns approved public or authorized nonprofit and IRS Form 990 sources into a structured dataset matched to your fields, filters, and delivery workflow. Review a sample before broader collection.
- Approved sources only
- Custom schema
- Flexible delivery

Nonprofit research becomes unreliable at dataset scale
Grant, diligence, and research teams often assemble nonprofit context from one-off EIN lookups, mismatched filing years, and naming variations that make identity matching and historical comparison harder as list volume grows.
Many filings arrive as PDF-only or document-only packages, and optional fields may be blank when the source does not display them. Treating missing values as collection failures creates false confidence in incomplete rows.
Internal scrapers and ad-hoc scripts add maintenance overhead when layouts, filters, or access rules change. At dataset scale, teams need a managed workflow that keeps source boundaries, validation notes, and delivery rules visible rather than hidden behind silent defaults.
What the GuideStar / ProPublica Nonprofit Scraper Service Provides
Nenodata scopes managed collection around approved public ProPublica nonprofit and IRS Form 990 records when those sources are in scope, and around authorized Candid/GuideStar access when the customer has confirmed permitted credentials or licensed access for the intended use.
Public ProPublica records and authorized Candid/GuideStar access are treated as distinct source lanes. Fields available on one lane are not assumed to exist on the other, and login-protected or subscription interfaces are not collected without explicit authorization and scoping.
Engagements map agreed organization identity, classification, filing metadata, financial fields where present, source URLs, collection timestamps, and validation status into a schema matched to your filters and destination. Broader managed programs may extend through Nenodata fully managed web scraping services when multi-source public-data collection is in scope. Source sets, fields, cadence, and destinations are agreed during scoping.
Illustrative sample output
Review an illustrative table and JSON record with organization identity, EIN, filing year, financial placeholders, source URL, collection timestamp, and validation status before broader production begins.
Illustrative example.

| organization_name | ein | filing_year | total_revenue | total_expenses | source_url | collected_at | validation_status |
|---|---|---|---|---|---|---|---|
| <Organization Name> | XX-XXXXXXX | YYYY | null | null | https://approved-source.example/nonprofit/XX-XXXXXXX/YYYY | YYYY-MM-DDTHH:mm:ssZ | pass_with_exceptions |
| <Organization Name B> | XX-XXXXXXY | YYYY | null | null | https://approved-source.example/nonprofit/XX-XXXXXXY/YYYY | YYYY-MM-DDTHH:mm:ssZ | pass_with_exceptions |
{
"organization_name": "<Organization Name>",
"ein": "XX-XXXXXXX",
"filing_year": "YYYY",
"total_revenue": null,
"total_expenses": null,
"source_url": "https://approved-source.example/nonprofit/XX-XXXXXXX/YYYY",
"collected_at": "YYYY-MM-DDTHH:mm:ssZ",
"validation_status": "pass_with_exceptions"
}Blank or unavailable values indicate that a field was not present in the approved source or was outside the agreed schema. They are not treated as collection failures unless validation rules mark them as errors.
Nonprofit data fields and delivery outputs
Field groups below are candidates for scoping. Availability depends on the approved source lane and the agreed schema.
Organization identity
Organization name, EIN, and related public identity labels when present on the approved source and included in scope.
Classification
IRS or NTEE-style classification labels when displayed and requested for the engagement—not assumed for every record.
Filing metadata
Filing year, form type, and related filing labels when scoped and available from the approved source method.
Financial information
Revenue, expense, and related financial fields only when present on the approved source and included in the agreed schema.
Traceability and quality
Source URLs, collection timestamps, validation status, and exception notes so blank or unavailable values remain visible.
Delivery options
CSV, Excel, JSON, API-ready structures, scheduled files, and database or warehouse destinations where technically feasible and confirmed.
Use cases
Grant prospect research
Teams assembling prospect lists struggle to keep EINs, filing years, and organization names aligned across sources. Structured records with agreed identity and filing fields support reviewable prospect libraries.
Nonprofit screening and due diligence
Diligence workflows stall when analysts re-key Form 990 observations into spreadsheets. Scoped extracts with source URLs and validation status support consistent screening packets.
Financial benchmarking
Benchmarking fails when revenue and expense fields are incomplete or inconsistently labeled. Normalized financial fields—when present—support internal comparisons without inventing missing values.
EIN-based enrichment
CRM and research systems often hold EINs without current filing context. EIN-keyed enrichment can attach agreed public or authorized fields when those values are available and in scope.
Filing monitoring
Recurring monitoring needs a defined cadence and exception handling rather than one-off downloads. Scheduled collection can refresh scoped organizations when refresh rules are contracted.
Journalism and public-interest research
Investigative teams need traceable public observations rather than opaque bulk dumps. Structured records with source references support public-interest research libraries within approved use limits.
Who this service is for
This service is for grantmakers, diligence teams, data vendors, research groups, journalists, and internal data engineering teams that need structured nonprofit and IRS Form 990 records matched to agreed fields and delivery destinations.
It fits organizations that can confirm source availability—public ProPublica records where feasible, or authorized Candid/GuideStar access when credentials or licensed access are approved for the intended use.
It is not positioned for unrestricted subscription scraping, guaranteed complete historical coverage, or buyers seeking legal, tax, or compliance advice.
How the engagement works
The delivery pattern aligns with how Nenodata works across managed public-data and authorized-source engagements.

- Step 1
Share your requirements
Provide representative organizations or EIN lists, required fields, filters, intended use, refresh need, and preferred delivery destination.
- Step 2
Review sources and access
Nenodata separates public ProPublica collection from authorized Candid/GuideStar access, confirms feasibility, and prepares a representative sample before production scale.
- Step 3
Collect and structure
Approved records are collected and mapped into the agreed schema with normalization rules that keep missing values and exceptions visible.
- Step 4
Validate and deliver
Validated outputs are delivered once or on a recurring schedule through formats and destinations confirmed during scoping, with maintenance where included.
Why choose Nenodata
Clear source boundaries
Public ProPublica records and authorized Candid/GuideStar access are scoped as distinct lanes so unsupported access is not implied.
Sample-first scoping
Representative fields, filters, and volume are reviewed through a sample before broader collection commitments.
Requirements-led schemas
Field names, validation labels, and destination mapping are planned around the systems that will consume the data.
Visible exceptions and traceability
Source URLs, collection timestamps, and validation status stay with records so blank or questionable values remain reviewable.
Managed operational ownership
When included in scope, Nenodata maintains agreed collection, validation, and delivery handling rather than shifting every source change to internal engineering.
Delivery into your existing workflow
Delivery formats and destinations are agreed during scoping. Options remain conditional on technical feasibility and are not presented as automatically available for every engagement.

- CSV or Excel
- JSON or API-ready structures
- Scheduled files
- Database or warehouse delivery where technically feasible
- Webhooks or CRM destinations where confirmed
Recurring transformation and destination routing may extend through Nenodata custom data pipelines. Review API delivery options and plans and custom pricing before confirming scope. Named connectors are not promised without destination confirmation.
Frequently Asked Questions
Request a representative nonprofit data sample
Share representative organizations or EIN lists, required fields, source preferences, intended use, one-time or recurring need, and preferred output destination so Nenodata can scope the next step.
Include example records, field requirements, cadence, destination, and intended use when you contact Nenodata through the contact flow.