Platform
Feature list (SDK spec)
The capability checklist for the ScrapeFlow scrapeflow npm and PyPI libraries. Both ship from the same API spec, so features stay at parity.
Data APIs (client resources)
Each maps 1:1 to an SDK method in both libraries.
| web.scrape() | URL → Markdown, HTML, rawHtml, links, screenshot, or structured JSON. |
| web.search() | Web search with ranked results; optional inline scraping. |
| web.answers() | Multi-step research → sourced, schema-shaped answers. |
| web.crawl() | Async full-site crawl on SQS workers; poll or webhook. |
| web.map() | Discover every URL on a domain, with titles/descriptions. |
| brand.get() | Domain → typed company profile, logo, colors, design tokens. |
| brand.styleguide() | Extract colors, typography, spacing, shadows, components. |
| batches.create() | Up to 25,000 URLs in one SNS→SQS fan-out job. |
| monitors.create() | Scheduled change detection (EventBridge) with diffs. |
| workflows.run() | Chain scrape → enrich → deliver pipelines. |
| jobs.get() / jobs.wait() | Poll or await any async job to completion. |
| usage.get() / cost.get() | Credits, spend, and budget status for dashboards. |
Client features (both SDKs)
Parity between npm and PyPI is a release requirement.
| Bearer auth | Key via constructor or SCRAPEFLOW_API_KEY env var. |
| Configurable base URL | Point at prod, staging, or a self-hosted gateway. |
| Automatic retries | Exponential backoff on 429 / 5xx, configurable maxRetries. |
| Timeouts | Per-request and global timeout controls. |
| Typed models | TypeScript types + Python type hints / Pydantic models. |
| Typed errors | ScrapeFlowError, RateLimitError, AuthError, BudgetExceededError. |
| Pagination helpers | Auto-iterate large result sets (batches, map, crawl). |
| Streaming | for-await (JS) / generators (Python) for batch results. |
| Idempotency keys | Safe retries for create operations. |
| Webhook verification | Helper to validate signed monitor/crawl callbacks. |
| Async support | Native async client in Python; promise-based in JS. |
| Structured logging hooks | Pluggable request/response logging. |
Developer experience
What makes teams pick and stay on the SDK.
| Zero-config quickstart | Works with just a key; sensible defaults. |
| Framework recipes | Next.js, Express, FastAPI, Django, LangChain, LlamaIndex. |
| pandas / DataFrame export | batches.to_dataframe() in Python. |
| Vector-store adapters | One-call push to Pinecone, pgvector, Chroma. |
| CLI | scrapeflow scrape <url> for quick terminal use. |
| MCP server | Expose all resources as tools to coding agents. |
| Semantic versioning | Generated from lib/apiSpec.js — spec is source of truth. |
Build order
- 1. Generate typed clients from
lib/apiSpec.js(JS + Python). - 2. Ship
scrape,search,brand(sync). - 3. Add async jobs:
crawl,batches,monitors. - 4. Add retries, typed errors, streaming, pagination.
- 5. Publish to npm + PyPI, then the CLI and MCP server.