1. The Web-Scale Ingestion Challenge in Enterprise RAG
Building Retrieval-Augmented Generation (RAG) pipelines across public technical documentation, corporate knowledge bases, and regulatory sites requires a web crawler that can scale horizontally across multiple machines while strictly enforcing web politeness, honoring robots.txt, and detecting removed pages.
Single-process crawlers quickly become a bottleneck or risk hitting IP rate limits. Apache StormCrawler is the industry-standard SDK for building low-latency, scalable web crawlers atop Apache Storm. Today, we are excited to release native OpenCrawling support for Apache StormCrawler, enabling automated web crawling directly into your vector search engines.
2. Two Complementary Modules: Connector & Bolt
To seamlessly integrate Apache Storm's distributed stream processing model with OpenCrawling's event-driven microservices architecture, the integration is split into two specialized modules:
-
oc-stormcrawler-repository-connector: A Spring Boot repository connector component implementing OpenCrawling'sRepositoryConnectorSPI. It connects to the Apache Storm Nimbus REST API, monitors cluster status, launches or pauses crawl topologies, validates seed URLs, and provides web politeness and URL filtering. -
oc-stormcrawler-bolt: An Apache Storm bolt (OpenCrawlingBolt extends BaseRichBolt) that is deployed inside the distributed Storm worker processes. It intercepts parsed HTML tuples and status deletion signals directly from Storm streams and transforms them into standard Open Ingestion Standard (OIS) payloads.
Java 25 & Apache Storm 3.1.0 Ready: Both modules run on Temurin JDK 25 LTS. With Apache Storm upgraded to 3.1.0, JEP 486 Security Manager removal is natively addressed while fully supporting Java 25 bytecode and Virtual Threads.
3. Dual-Stream Processing & OIS Deletion Tombstones
Keeping vector databases synchronized with live websites requires not only indexing new and updated pages, but also purging stale vectors when documents are removed. OpenCrawlingBolt subscribes to two distinct streams:
-
Default Content Stream: Receives parsed documents from
ParserBoltcontaining(url, content, metadata). The bolt computes cryptographic content hashes (SHA-256orMD5) to eliminate duplicate vector embeddings and generates standard OISUPSERTrecords. -
Status Stream (
Constants.StatusStreamName): Receives crawl status tuples from fetcher components. When a URL returns HTTP 404 (Not Found), HTTP 410 (Gone), orFETCH_ERROR, the bolt emits an OISDELETEtombstone payload, instructing downstream vector store consumers to purge all associated chunks.
Zero Stale Vectors: When documentation pages are removed from target sites, OpenCrawling automatically purges vector chunks from PGVector, Milvus, Solr, Luxir, or Vespa, preventing LLM hallucinations on outdated knowledge.
4. Java 25 Structured Concurrency & Politeness
The StormCrawlerRepositoryConnector orchestrates seed URL resolution and pre-flight HTTP validations using Java 25 Virtual Threads and StructuredTaskScope. This enables high-concurrency URL discovery without exhausting OS thread pools.
Crawl politeness is governed by configurable delays (delayMs), custom user-agent headers (OpenCrawling-StormCrawler-Bot/1.0), robots.txt evaluation, and regular expression filters that prevent crawling binary assets like archives and executables.
5. Configuring the StormCrawler Connector
The connector is configured using standard Spring Boot properties in application.yml or environment variables:
spring:
opencrawling:
connector:
type: stormcrawler
stormcrawler:
nimbus-host: "storm-nimbus"
nimbus-port: 6627
nimbus-rest-url: "http://storm-ui:8080"
topology-name: "opencrawling-web-crawler"
concurrency: 8
delay-ms: 1000
seeds:
- "https://docs.opencrawling.org"
- "https://developer.opencrawling.org"
custom-user-agent: "OpenCrawling-StormCrawler-Bot/1.0"
exclude-patterns:
- ".*\\.(pdf|zip|gz|exe)$"
| Property | Environment Variable | Default | Description |
|---|---|---|---|
nimbus-host |
STORMCRAWLER_NIMBUS_HOST |
localhost |
Apache Storm Nimbus master hostname |
nimbus-port |
STORMCRAWLER_NIMBUS_PORT |
6627 |
Nimbus Thrift RPC service port |
nimbus-rest-url |
STORMCRAWLER_NIMBUS_REST_URL |
http://localhost:8080 |
Storm UI / Nimbus REST API base URL |
topology-name |
STORMCRAWLER_TOPOLOGY_NAME |
opencrawling-web-crawler |
Target StormCrawler crawl topology name |
delay-ms |
STORMCRAWLER_DELAY_MS |
1000 |
Politeness delay between requests (ms) |
concurrency |
STORMCRAWLER_CONCURRENCY |
8 |
Parallel virtual threads for crawl tasks |
6. Registering OpenCrawlingBolt in Storm Topologies
To attach OpenCrawling to any existing StormCrawler topology, include the oc-stormcrawler-bolt dependency and register the bolt using Storm's standard TopologyBuilder:
TopologyBuilder builder = new TopologyBuilder();
// 1. Configure OpenCrawling Bolt
OpenCrawlingBolt openCrawlingBolt = new OpenCrawlingBolt()
.withTargetEndpoint("http://oc-runtime:8080/api/v1/ingest/ois")
.withTransportMode("REST")
.withEmitDeletions(true)
.withInstanceId("storm-cluster-production");
// 2. Wire content and status deletion streams
builder.setBolt("opencrawling-bolt", openCrawlingBolt, 2)
.shuffleGrouping("parser") // Default content stream (UPSERT)
.shuffleGrouping("fetcher", Constants.StatusStreamName); // Status 404/410 stream (DELETE)
7. Docker Compose & Automated Integration Testing
The module comes with ready-to-run Docker Compose stacks for spinning up a distributed Apache Storm 3.1.0 cluster (ZooKeeper, Nimbus, Supervisor, UI) alongside the full OpenCrawling decoupled microservices pipeline.
You can execute the dedicated end-to-end integration test right from the terminal:
# Run dedicated StormCrawler Bolt cluster integration test
chmod +x scripts/test-stormcrawler-bolt.sh
./scripts/test-stormcrawler-bolt.sh
The script builds the shaded topology fat JAR, launches the distributed Storm cluster, deploys OpenCrawlingTestTopology to Nimbus, and verifies that opencrawling-bolt is actively processing crawl tuples via the Storm REST API.
Ready to Scale Your Web Crawling for Enterprise AI?
Explore the complete documentation on the OpenCrawling Wiki and start deploying distributed web crawl pipelines today.