The OpenCrawling ecosystem continues to expand rapidly as we bridge enterprise document repositories, relational databases, content systems, and web data sources to modern AI vector databases. As our architectural footprint widens, so does our community of passionate engineers.
Today, I am thrilled to officially announce that Davide Polato (@dpol1) has joined the OpenCrawling project as a Contributor for Apache Storm & StormCrawler.
Distributed Stream Processing & Web Crawling
Davide joins to reinforce OpenCrawling's distributed web crawling capabilities, bringing deep hands-on expertise in Apache Storm topologies, StormCrawler components, and test suite determinism.
Distributed Web Crawling for Enterprise AI
While enterprise content systems (Alfresco, SharePoint, CMIS, Documentum, FileNet) and structured databases host proprietary internal knowledge, modern enterprise RAG and LLM agent systems frequently require web-scale crawling of documentation hubs, product sites, knowledge portals, and authenticated partner domains.
To address this need, OpenCrawling recently introduced native support for Apache StormCrawler through two key modules:
-
oc-stormcrawler-bolt: An Apache Storm Bolt that intercepts parsed web pages and emits normalized Open Ingestion Standard (OIS) documents with cryptographic deduplication, as well as HTTP 404/410 lifecycle events as OIS deletion tombstones (action: "DELETE"). -
oc-stormcrawler-repository-connector: An OpenCrawling repository connector that coordinates StormCrawler topologies via the Apache Storm Nimbus REST API, handling lifecycle supervision and scheduling directly from the OpenCrawling Admin UI.
Running large-scale distributed stream topologies across Apache Storm clusters introduces unique challenges: asynchronous tuple emission, frontier synchronization, stateful URL tracking, and testing non-determinism under parallel execution. This is where Davide's expertise shines.
Engineering Rigor & Test Determinism
Davide is a software engineer and computer science researcher focused on Java backend internals, distributed stream processing, and clean code principles. As an active open-source contributor to both Apache Storm and Apache StormCrawler, he has been deeply engaged in solving complex distributed systems problems, including charset resolution, distributed mode stability, and observability context propagation.
Davide recently made an invaluable contribution to OpenCrawling by resolving a tricky, non-deterministic test issue in our StormCrawler repository connector test suite. In parallel stream topologies, tuples can be emitted in non-sequential order across worker threads. Davide collaborated with the core team to ensure unit and integration tests assert scan results independently of emission order, dramatically improving the determinism and reliability of our continuous integration pipelines.
"Distributed stream processing with Apache Storm and StormCrawler offers extraordinary throughput and scalability for web crawling. Bringing that stream-first power into OpenCrawling's Open Ingestion Standard—while ensuring bulletproof test reliability and clean code architecture—is an exciting journey. I am delighted to collaborate with the team to push the boundaries of distributed enterprise data ingestion."
— Davide Polato, Contributor for Apache Storm & StormCrawler
What Davide Will Be Focusing on Next
Davide is actively collaborating on the OpenCrawling codebase. His immediate roadmap includes:
- Deepening Storm & StormCrawler Integration: Ensuring complete harmony between Apache Storm 3.1.0, StormCrawler 3.7.0, and the Open Ingestion Standard (OIS) document model.
- Distributed Test Reliability: Enhancing test harnesses for distributed crawler topologies, status stream emission, and tombstone lifecycle events to maintain high quality and CI determinism.
- Performance & Frontier Tuning: Optimizing low-latency streaming pipeline execution, backpressure management, and URL frontier coordination.
- Cross-Community Synergy: Strengthening ties and collaboration between the Apache Storm, Apache StormCrawler, and OpenCrawling developer communities.
A Note from Piergiorgio Lucidi
“Building a truly robust open-source ingestion platform requires not only visionary architectural design but also rigorous engineering discipline under the hood. Davide's mastery of Apache Storm, his deep understanding of distributed stream mechanics, and his commitment to clean, maintainable code make him a fantastic addition to our team. Having him join as a Contributor for Apache Storm & StormCrawler ensures that our web crawling capabilities are second to none in reliability and throughput. Welcome to the team, Davide!”
Join Us in Welcoming Davide!
You can find Davide's profile on our Team Page, follow his open-source work on GitHub (@dpol1), connect with him on LinkedIn, or chat with him directly in the OpenCrawling Slack Community!