Connector Release September 28, 2026 • 6 min read

Enterprise Distributed Web Crawling: Announcing the Apache StormCrawler Repository Connector & OpenCrawling Bolt

OpenCrawling expands its repository ecosystem to web-scale crawling. Built with Apache StormCrawler 3.7.0, Apache Storm 3.1.0, and Java 25, the new oc-stormcrawler-repository-connector and oc-stormcrawler-bolt deliver distributed stream ingestion, real-time HTTP 404/410 deletion tombstones, and Nimbus REST API supervision for production AI pipelines.

Distributed Web Crawling Pipeline Flow
Web Seed URLs
→
StormCrawler Topology
→
OpenCrawlingBolt
→
OIS REST Ingestion Queue
→
Apache Kafka
→
Vector Stores

1. The Web-Scale Ingestion Challenge in Enterprise RAG

Building Retrieval-Augmented Generation (RAG) pipelines across public technical documentation, corporate knowledge bases, and regulatory sites requires a web crawler that can scale horizontally across multiple machines while strictly enforcing web politeness, honoring robots.txt, and detecting removed pages.

Single-process crawlers quickly become a bottleneck or risk hitting IP rate limits. Apache StormCrawler is the industry-standard SDK for building low-latency, scalable web crawlers atop Apache Storm. Today, we are excited to release native OpenCrawling support for Apache StormCrawler, enabling automated web crawling directly into your vector search engines.

2. Two Complementary Modules: Connector & Bolt

To seamlessly integrate Apache Storm's distributed stream processing model with OpenCrawling's event-driven microservices architecture, the integration is split into two specialized modules:

ℹ️

Java 25 & Apache Storm 3.1.0 Ready: Both modules run on Temurin JDK 25 LTS. With Apache Storm upgraded to 3.1.0, JEP 486 Security Manager removal is natively addressed while fully supporting Java 25 bytecode and Virtual Threads.

3. Dual-Stream Processing & OIS Deletion Tombstones

Keeping vector databases synchronized with live websites requires not only indexing new and updated pages, but also purging stale vectors when documents are removed. OpenCrawlingBolt subscribes to two distinct streams:

  1. Default Content Stream: Receives parsed documents from ParserBolt containing (url, content, metadata). The bolt computes cryptographic content hashes (SHA-256 or MD5) to eliminate duplicate vector embeddings and generates standard OIS UPSERT records.
  2. Status Stream (Constants.StatusStreamName): Receives crawl status tuples from fetcher components. When a URL returns HTTP 404 (Not Found), HTTP 410 (Gone), or FETCH_ERROR, the bolt emits an OIS DELETE tombstone payload, instructing downstream vector store consumers to purge all associated chunks.
✅

Zero Stale Vectors: When documentation pages are removed from target sites, OpenCrawling automatically purges vector chunks from PGVector, Milvus, Solr, Luxir, or Vespa, preventing LLM hallucinations on outdated knowledge.

4. Java 25 Structured Concurrency & Politeness

The StormCrawlerRepositoryConnector orchestrates seed URL resolution and pre-flight HTTP validations using Java 25 Virtual Threads and StructuredTaskScope. This enables high-concurrency URL discovery without exhausting OS thread pools.

Crawl politeness is governed by configurable delays (delayMs), custom user-agent headers (OpenCrawling-StormCrawler-Bot/1.0), robots.txt evaluation, and regular expression filters that prevent crawling binary assets like archives and executables.

5. Configuring the StormCrawler Connector

The connector is configured using standard Spring Boot properties in application.yml or environment variables:

spring:
  opencrawling:
    connector:
      type: stormcrawler
      stormcrawler:
        nimbus-host: "storm-nimbus"
        nimbus-port: 6627
        nimbus-rest-url: "http://storm-ui:8080"
        topology-name: "opencrawling-web-crawler"
        concurrency: 8
        delay-ms: 1000
        seeds:
          - "https://docs.opencrawling.org"
          - "https://developer.opencrawling.org"
        custom-user-agent: "OpenCrawling-StormCrawler-Bot/1.0"
        exclude-patterns:
          - ".*\\.(pdf|zip|gz|exe)$"
Property Environment Variable Default Description
nimbus-host STORMCRAWLER_NIMBUS_HOST localhost Apache Storm Nimbus master hostname
nimbus-port STORMCRAWLER_NIMBUS_PORT 6627 Nimbus Thrift RPC service port
nimbus-rest-url STORMCRAWLER_NIMBUS_REST_URL http://localhost:8080 Storm UI / Nimbus REST API base URL
topology-name STORMCRAWLER_TOPOLOGY_NAME opencrawling-web-crawler Target StormCrawler crawl topology name
delay-ms STORMCRAWLER_DELAY_MS 1000 Politeness delay between requests (ms)
concurrency STORMCRAWLER_CONCURRENCY 8 Parallel virtual threads for crawl tasks

6. Registering OpenCrawlingBolt in Storm Topologies

To attach OpenCrawling to any existing StormCrawler topology, include the oc-stormcrawler-bolt dependency and register the bolt using Storm's standard TopologyBuilder:

TopologyBuilder builder = new TopologyBuilder();

// 1. Configure OpenCrawling Bolt
OpenCrawlingBolt openCrawlingBolt = new OpenCrawlingBolt()
    .withTargetEndpoint("http://oc-runtime:8080/api/v1/ingest/ois")
    .withTransportMode("REST")
    .withEmitDeletions(true)
    .withInstanceId("storm-cluster-production");

// 2. Wire content and status deletion streams
builder.setBolt("opencrawling-bolt", openCrawlingBolt, 2)
    .shuffleGrouping("parser")                                  // Default content stream (UPSERT)
    .shuffleGrouping("fetcher", Constants.StatusStreamName);    // Status 404/410 stream (DELETE)

7. Docker Compose & Automated Integration Testing

The module comes with ready-to-run Docker Compose stacks for spinning up a distributed Apache Storm 3.1.0 cluster (ZooKeeper, Nimbus, Supervisor, UI) alongside the full OpenCrawling decoupled microservices pipeline.

You can execute the dedicated end-to-end integration test right from the terminal:

# Run dedicated StormCrawler Bolt cluster integration test
chmod +x scripts/test-stormcrawler-bolt.sh
./scripts/test-stormcrawler-bolt.sh

The script builds the shaded topology fat JAR, launches the distributed Storm cluster, deploys OpenCrawlingTestTopology to Nimbus, and verifies that opencrawling-bolt is actively processing crawl tuples via the Storm REST API.

Ready to Scale Your Web Crawling for Enterprise AI?

Explore the complete documentation on the OpenCrawling Wiki and start deploying distributed web crawl pipelines today.