Platform

Feature list (SDK spec)

The capability checklist for the ScrapeFlow scrapeflow npm and PyPI libraries. Both ship from the same API spec, so features stay at parity.

Data APIs (client resources)

Each maps 1:1 to an SDK method in both libraries.

web.scrape()URL → Markdown, HTML, rawHtml, links, screenshot, or structured JSON.
web.search()Web search with ranked results; optional inline scraping.
web.answers()Multi-step research → sourced, schema-shaped answers.
web.crawl()Async full-site crawl on SQS workers; poll or webhook.
web.map()Discover every URL on a domain, with titles/descriptions.
brand.get()Domain → typed company profile, logo, colors, design tokens.
brand.styleguide()Extract colors, typography, spacing, shadows, components.
batches.create()Up to 25,000 URLs in one SNS→SQS fan-out job.
monitors.create()Scheduled change detection (EventBridge) with diffs.
workflows.run()Chain scrape → enrich → deliver pipelines.
jobs.get() / jobs.wait()Poll or await any async job to completion.
usage.get() / cost.get()Credits, spend, and budget status for dashboards.

Client features (both SDKs)

Parity between npm and PyPI is a release requirement.

Bearer authKey via constructor or SCRAPEFLOW_API_KEY env var.
Configurable base URLPoint at prod, staging, or a self-hosted gateway.
Automatic retriesExponential backoff on 429 / 5xx, configurable maxRetries.
TimeoutsPer-request and global timeout controls.
Typed modelsTypeScript types + Python type hints / Pydantic models.
Typed errorsScrapeFlowError, RateLimitError, AuthError, BudgetExceededError.
Pagination helpersAuto-iterate large result sets (batches, map, crawl).
Streamingfor-await (JS) / generators (Python) for batch results.
Idempotency keysSafe retries for create operations.
Webhook verificationHelper to validate signed monitor/crawl callbacks.
Async supportNative async client in Python; promise-based in JS.
Structured logging hooksPluggable request/response logging.

Developer experience

What makes teams pick and stay on the SDK.

Zero-config quickstartWorks with just a key; sensible defaults.
Framework recipesNext.js, Express, FastAPI, Django, LangChain, LlamaIndex.
pandas / DataFrame exportbatches.to_dataframe() in Python.
Vector-store adaptersOne-call push to Pinecone, pgvector, Chroma.
CLIscrapeflow scrape <url> for quick terminal use.
MCP serverExpose all resources as tools to coding agents.
Semantic versioningGenerated from lib/apiSpec.js — spec is source of truth.

Build order

  1. 1. Generate typed clients from lib/apiSpec.js (JS + Python).
  2. 2. Ship scrape, search, brand (sync).
  3. 3. Add async jobs: crawl, batches, monitors.
  4. 4. Add retries, typed errors, streaming, pagination.
  5. 5. Publish to npm + PyPI, then the CLI and MCP server.