New Release
August 7, 2026
6 min read
Discover OpenCrawling's native Vespa Output Connector (oc-vespa-output-connector). Built with Yahoo's HTTP/2 vespa-feed-client, dynamic multi-dimension vector routing (384/768/1024), ACL pre-filtering, Admin UI Model Insights, and Vespa-native Secure MCP tools.
Vector Store
July 29, 2026
5 min read
Discover OpenCrawling's native Qdrant Output Connector (oc-qdrant-output-connector). Built with Qdrant's high-performance gRPC Java SDK, scalar quantization, and automatic payload ACL indexing for secure, low-latency enterprise RAG.
SDK / Java 25
July 28, 2026
5 min read
Discover the official oc-java-client-sdk. Strongly typed Java 25 & Spring Boot 4 client library for programmatically managing document ingestion jobs, connectors, Auto-Narrativization Copilot templates, and AIOps diagnostics.
Maven / Archetypes
July 26, 2026
4 min read
Learn how OpenCrawling's official Maven Archetypes suite allows developers to instantly scaffold ready-to-build Repository, Output, and Transformation Connectors with Docker Compose overlays and Admin UI integration.
Workflow / BPMN
July 24, 2026
6 min read
OpenCrawling now natively supports the top open-source BPMN workflow engines: Flowable and Camunda. Discover how oc-flowable-repository-connector and oc-camunda-repository-connector stream process definitions, historic execution variables, and candidate group ACLs into AI pipelines.
AI / Metadata
July 23, 2026
6 min read
Learn how OpenCrawling's Auto-Narrativization Copilot translates structured database schemas and table records into rich natural language Mustache templates and mock datasets using Spring AI, Ollama (llama3.2), and deterministic offline fallbacks.
Storage
July 21, 2026
6 min read
OpenCrawling now offers full, native support for Apache Ozone (v2.2.0+) as a distributed Claim Check store provider. Learn how combining the Claim Check Pattern with Ozone's S3 Gateway eliminates Kafka payload bottlenecks and scales RAG binary ingestion to petabyte scale.
Connector
July 16, 2026
5 min read
OpenCrawling now supports Apache Iceberg data lakes as a first-class ingestion source. The new oc-iceberg-repository-connector module scans Iceberg table catalogs (REST, Hive, Hadoop, AWS Glue) and streams structured records directly into the RAG pipeline.
Connector
July 16, 2026
6 min read
The new oc-alfresco-repository-connector bridges Alfresco Content Services to your AI pipeline via the Alfresco REST API v1. Discover how Java 25 Structured Task Scope enables parallel folder traversal, and how the Claim Check Pattern decouples remote content downloads from the Kafka publish step.
Team
July 16, 2026
3 min read
Enterprise Content Management (ECM) integration specialist Luis Cabaceira joins the OpenCrawling core team. Read about his work on repository connectors, CMIS mappings, and secure enterprise document ingestion.
Team
July 14, 2026
4 min read
Enterprise search architect and big data veteran Michael Cizmar joins the OpenCrawling core team. Read about his background and his plans to drive event-driven ingestion parallelization and multi-tenant vector writer integrations.
Architecture
July 12, 2026
7 min read
Discover how the new oc-embedding-service microservice separates embedding generation into a dedicated, horizontally scalable unit. Learn how Kafka consumer group partitioning, the EmbeddingModelFactory, and Docker Compose scaling unlock GPU-level throughput without changing the rest of your pipeline.
Technology
July 10, 2026
6 min read
Learn how OpenCrawling leverages the Model Context Protocol (MCP) and Spring AI to expose secure, enterprise-grade, ACL-filtered knowledge retrieval tools directly to LLM agents.
Announcement
July 9, 2026
5 min read
Enterprise data pipelines suffer from proprietary lock-in and security leaks. Today we are launching OIS—a Zero-Trust, vendor-neutral specification—and its Java 25 reference engine, OpenCrawling. Learn how we preserve ACL security.
Roadmap
July 9, 2026
4 min read
A security mapping framework is only as powerful as the repositories it can reach. Share your feedback in our GitHub Roadmap discussions and vote on SharePoint, S3, Qdrant, and Confluence.