# CrawlKit Docs CrawlKit is an Arrow-native web data engineering and intelligence platform: crawling, spider management, scraping, ETL, DataFusion/DuckDB query, datasets, durable workflows, quality, lineage, enrichment, OpenRate healthcare pricing, SEO/GSC, telemetry, and a canonical 263-tool registry shared by Pi and the universal MCP bridge. Canonical URLs: - API base: https://api.crawlkit.app/api/v1 - Docs: https://docs.crawlkit.app/ - Platform model: https://docs.crawlkit.app/platform.html - OpenAPI: https://docs.crawlkit.app/openapi.json - Complete API surface: https://docs.crawlkit.app/capabilities.html - Arrow/DataFusion guide: https://docs.crawlkit.app/guides/data-engineering.html - Workflows guide: https://docs.crawlkit.app/guides/workflows.html - Spiders guide: https://docs.crawlkit.app/guides/spiders.html - Google Search Console setup: https://docs.crawlkit.app/guides/google-search-console-setup.html - Google Search Console usage: https://docs.crawlkit.app/guides/google-search-console.html - Reusable authorization pattern: https://docs.crawlkit.app/guides/authorization-integrations.html - Full-parity agent integrations (Hermes, Claude, Codex, OpenCode): https://docs.crawlkit.app/guides/integrations.html - Tutorials index: https://docs.crawlkit.app/tutorials/ - Full endpoint inventory: https://docs.crawlkit.app/llms-full.txt Current public API contract: 383 operations, 315 paths, 54 tags, 532 schemas. Priority domains for technical users: Query, ETL, Workflow Templates, Pipelines, Spiders, Crawl, Datasets, Artifacts, Lineage, Quality, SchemaRegistry, Enrichment, OpenRate, Telemetry, Intelligence. Canonical platform narrative: CrawlKit is a data engineering backbone plus capability workbenches plus downstream applications. The operating loop is Acquire -> Extract -> Refine -> Serve -> Prove. Tutorials remain available and should be interpreted as core platform tutorials or applied lenses over the same backbone.