The momentum behind OpenCrawling continues to accelerate. As we expand the reach of the Open Ingestion Standard (OIS) to encompass complex enterprise file structures and workflows, we are also growing the developer network building the core connectors.

Today, I am absolutely delighted to announce that Enterprise Content Management (ECM) integration expert Luis Cabaceira has officially joined the OpenCrawling core steering group as a Core Contributor and Lead Architect.

🗂️

Enterprise Ingestion Connectors

Luis joins to lead repository connector development, CMIS mappings, and security-aware content extraction from legacy ECM repositories.

Unlocking Legacy Enterprise Content (ECM) for AI

For an open standard like OIS to thrive, it must securely bridge modern vector databases with legacy repository infrastructures. Enterprise content management platforms store the core intellectual property of organizations, but extracting it securely has always been a major bottleneck. If you bypass ACL checks or strip metadata, the vectorized data becomes both a compliance hazard and contextually poor.

Luis brings an outstanding background in enterprise search logistics and document repositories. His professional experience in developing CMIS integrations, SharePoint hooks, and document indexing pipelines makes him uniquely qualified to shape OpenCrawling’s connector framework. Luis will be steering the development of new out-of-the-box repository connectors, ensuring legacy systems can feed unstructured assets into secure AI pipelines in minutes.

"Enterprise content management platforms hold the core knowledge of an organization, but indexing them for RAG while preserving complex security permissions has historically been a challenge. I am excited to help expand OpenCrawling's repository connector suite, making it simple to bridge legacy ECM systems with modern vector stores."

— Luis Cabaceira, Core Contributor and Lead Architect

What Luis is focus-structuring next

Luis is already active in our codebase. His immediate technical focus will be centered on:

Coming Next: The Vespa Output Connector

Alongside his repository connector work, Luis is spearheading an exciting addition to the OpenCrawling output connector suite: the oc-vespa-output-connector. This module will allow the pipeline to write vectorized document chunks directly to Vespa, Yahoo’s battle-tested open-source search and recommendation engine, which combines ANN-based vector search with structured filtering, ranking expressions, and real-time updates in a single platform.

🚀

oc-vespa-output-connector — In Development

Luis is actively working on the Vespa output connector, which will join oc-pgvector-output-connector and the Elasticsearch connector as a first-class output target for the OpenCrawling ingestion pipeline. Follow the progress on GitHub or join the discussion in our Slack community.

The Vespa connector will implement the standard OutputConnector interface from oc-core, and will support both the direct output mode (single-JVM, no Kafka) and the decoupled Kafka pipeline, consuming from the opencrawling-embedded topic and writing to a Vespa application schema. Vespa’s unique combination of lexical, semantic, and structured query operators makes it an exceptionally powerful target for enterprise RAG systems that need ranked results beyond pure cosine similarity.

A Note from Piergiorgio Lucidi

“A security mapping framework is only as good as the systems it can securely connect to. Luis’s deep expertise in Enterprise Content Management (ECM) platforms and information logistics is exactly what we need. To build a world-class ingestion pipeline, we must be able to securely connect to where the data lives—and deliver results to where the AI needs them. Luis’s skills will accelerate our connector roadmap, both unlocking massive legacy repositories and expanding our vector store ecosystem with production-grade targets like Vespa.”

Join Us in Welcoming Luis!

We are building a highly collaborative, open-source community. You can find Luis's profile on our Team Page, track his commits on GitHub, or connect with him directly in our Slack Community!