Skip to main content
Diffbot logo

Diffbot

Automated web extraction and a trillion-fact Knowledge Graph

automation#web-scraping#knowledge-graph#data-extraction#api
Free plan Free trial Claimed API Teams
Toolglade’s take

Diffbot is unusually differentiated: its automatic, vision-based extraction and its Knowledge Graph of a trillion-plus facts are hard for most competitors to match, and it is a favorite among developers who need clean structured data or entity intelligence at scale. The trade-offs are a credit-based pricing model that can get expensive for Knowledge Graph queries, a developer-first experience that is not no-code, and, as with any large-scale crawl-derived dataset, the need to validate coverage and accuracy for your specific entities. It is a strong pick for technical teams, less so for non-technical users.

About Diffbot

Diffbot automatically extracts structured data from web pages using computer vision and ML, and maintains a Knowledge Graph of over 10 billion entities and a trillion-plus facts. It is a developer-first platform with extraction APIs, Crawlbot, and Enhance, offering a free credit tier and credit-based paid plans. It is a standout for web-scale structured data and entity intelligence, but is technical and can be costly at scale.

Diffbot's extraction APIs (for articles, products, discussions, and more) automatically parse arbitrary web pages into structured JSON using computer vision and ML, without site-specific rules. Its Crawlbot can crawl entire sites, and Diffbot Enhance can append firmographic and entity data to records. The result is that developers can extract clean structured data from pages they have never configured for. The flagship differentiator is the Diffbot Knowledge Graph, a machine-constructed graph reportedly spanning over 10 billion entities and more than a trillion structured facts, refreshed continuously from crawling the public web. It powers use cases in market intelligence, lead enrichment, news monitoring, and research, queried via API. Diffbot was founded in 2008 (formally around 2012) by Mike Tung and is based in the Bay Area. It raised a relatively modest amount of venture funding, including a 2016 Series A led by Tencent and Felicis, and has operated as a focused, technically deep company. It is developer-oriented and priced with a credit model that can suit both experimentation and heavy production use.

TL;DR

Diffbot automatically extracts structured data from any web page using computer vision and ML, and maintains a Knowledge Graph of over 10 billion entities and a trillion-plus facts. Founded in 2008 by Mike Tung, it raised a modest ~$12.5M including a 2016 Tencent/Felicis Series A. It is a developer-first platform with extraction APIs, Crawlbot, and Enhance, priced on credits with a free tier. Strong for web-scale data and entity intelligence.

Company overview

Diffbot was founded by Mike Tung (early work from 2008, company established in the early 2010s) and is based in the San Francisco Bay Area. It has stayed a focused, technically deep company centered on automatic web extraction and its Knowledge Graph.

Diffbot raised a relatively modest amount of venture capital, around $12.5M total, including a $10M Series A in 2016 led by Tencent and Felicis Ventures, with Bloomberg Beta among earlier backers. It has operated with a lean, engineering-driven model rather than aggressive fundraising.

Product features

Its extraction APIs turn arbitrary pages into structured JSON without site-specific rules, Crawlbot crawls entire sites, and Enhance appends entity and firmographic data. The Knowledge Graph, queryable via API, spans billions of entities and a trillion-plus facts refreshed from continuous crawling.

Together these enable market intelligence, enrichment, news monitoring, and research at scale, all delivered as clean structured data for developers to integrate.

Target market

Developers, data engineers, and data-driven teams needing automatic structured extraction or entity intelligence at web scale, including market intelligence, sales enrichment, and research use cases.

Buyer personas

End users

Developers and data engineers integrating extraction APIs and Knowledge Graph queries.

Buyers

Engineering, data, and product leaders funding web-data or enrichment infrastructure.

Key influencers

Technical evaluators, data-science teams, and market-intelligence practitioners.

Ideal customer profile

A technical team needing automatic, web-scale structured extraction or entity intelligence, comfortable with API-driven, credit-based tooling.

Funding & performance

Diffbot raised approximately $12.5M total, including a $10M Series A in 2016 led by Tencent and Felicis Ventures, with Bloomberg Beta among earlier investors. Figures per public reporting; verify with the company.

Pros & cons

Pros

  • Automatic, vision-based extraction without site rules
  • Large Knowledge Graph of billions of entities
  • Powerful Crawlbot and Enhance capabilities
  • Clean structured JSON output
  • Strong developer documentation and API
  • Free credit tier to evaluate

Cons

  • Credit-based pricing can get expensive at scale
  • Knowledge Graph queries consume many credits
  • Developer-first; not a no-code tool
  • Coverage and accuracy vary by entity and source
  • No self-hosted option
  • Overkill for simple single-site scraping

Pricing plans

Free
$0 / month
  • ~10,000 credits
  • Rate-limited API access
  • Dashboard access
  • For evaluation
Startup
Custom / month
  • ~250,000 credits
  • Higher rate limits
  • Full extraction API access
  • Standard support
Plus
Custom / month
  • ~1,000,000 credits
  • Higher throughput
  • Crawlbot access
  • Knowledge Graph access
Enterprise
Custom
  • Custom credits and limits
  • Dedicated success manager
  • Phone support
  • Tailored features

Key features

API
Team collaboration
Multi-language
Integrations
REST API, Crawlbot, Knowledge Graph API, Enhance API, Zapier, Excel/Sheets connectors
Input types
url, text, query
Output types
json, structured data, entity records
Best For
Automatic structured extraction from any page, Querying a large Knowledge Graph, Entity and firmographic enrichment, News and market intelligence at scale

Compare key features

View all alternatives →
Feature
Diffbot
Browse AI
Apify
Pricing
Freemium
Freemium
Freemium
Free plan
Yes
Yes
Yes
Free trial
Yes
No
Yes
API
Yes
Yes
Yes
Team support
Yes
Yes
Yes

Frequently asked questions

How does Diffbot extract data without site-specific rules?+

It uses computer vision and machine learning to recognize page structures (articles, products, discussions), so it can parse pages it has never seen into structured JSON automatically.

What is the Diffbot Knowledge Graph?+

It is a machine-built graph of over 10 billion entities and more than a trillion structured facts, continuously refreshed from crawling the public web and queryable via API.

Is there a free way to try Diffbot?+

Yes. Diffbot offers a free tier with a limited number of credits and rate limits so you can evaluate the extraction APIs and Knowledge Graph.

Why can Diffbot get expensive?+

Pricing is credit-based, and heavy usage, especially Knowledge Graph queries that cost many credits each, can add up quickly at scale.

Is Diffbot suitable for non-developers?+

Not really. It is a developer-first, API-driven platform. Non-technical users typically need engineering help or a no-code layer on top.

Reviews (0)

Write a review

Pick a rating
Loading reviews…
Compare

Compare Diffbot with other AI tools

Side-by-side pages for pricing, features, and best-fit use cases.

All comparisons →

Similar tools you may like