Connector Release October 1, 2026 • 7 min read

Universal ECM Ingestion for Enterprise RAG: Announcing the OASIS CMIS Repository Connector

Enterprise document ecosystems are famously fragmented across decades of vendor solutions: Alfresco Content Services, OpenText Documentum, IBM FileNet P8, Nuxeo, and Apache Chemistry repositories. Today, we are releasing oc-cmis-repository-connector, delivering universal, high-throughput ingestion of folders, CMIS-SQL queries, version series, aspects, and Zero-Trust Access Control Lists (ACLs) to OpenCrawling via the OASIS CMIS 1.0 & 1.1 open standard! 🎉

OASIS CMIS Ingestion Pipeline Architecture
ECM (Alfresco / FileNet / Documentum)
→
CmisClient (Java 25 Browser Binding)
→
Claim-Check Store (Ozone / S3)
→
OIS with Zero-Trust ACEs
→
Kafka Broker
→
Vector Stores (PGVector / Solr / Milvus)

1. The Case for an Open Standard in Enterprise AI

Building Retrieval-Augmented Generation (RAG) applications in Fortune 500 and public-sector environments rarely involves a single, clean content repository. Large enterprises often maintain historical contracts in IBM FileNet, regulated engineering designs in OpenText Documentum, collaborative project spaces in Alfresco Content Services, and departmental archives in Nuxeo.

Creating bespoke, vendor-specific crawlers for every proprietary ECM API leads to technical debt, mismatched metadata schemas, and fragile security boundaries.

The Content Management Interoperability Services (CMIS) specification—standardized by OASIS—was created to solve exactly this challenge. By adopting CMIS as a universal abstraction layer, OpenCrawling enables organizations to connect once and ingest from any compliant ECM system, preserving folder hierarchies, custom aspects, document versions, and fine-grained permissions without vendor lock-in.

2. Pure Java 25 & Zero Legacy Dependencies

Historically, Java integrations with CMIS relied heavily on the Apache Chemistry OpenCMIS library. While Apache Chemistry pioneered CMIS adoption, its client libraries carry dozens of legacy transitive dependencies, heavy XML binding frameworks (JAXB/JAX-WS), and synchronous I/O architectures ill-suited for modern reactive microservices and containerized runtimes.

OpenCrawling's oc-cmis-repository-connector is built with a fundamentally modern approach:

ℹ️

Zero Chemistry Footprint: By eliminating legacy Chemistry JARs, oc-cmis-repository-connector runs cleanly on modular Java 25 runtimes, consumes up to 70% less heap memory during massive scans, and supports instant container scaling.

3. Multi-Mode Crawling: Hierarchies, Queries & Change Logs

Different enterprise search and compliance use cases demand different ingestion strategies. The CMIS connector supports three distinct crawl modes:

  1. Recursive Folder Crawl (crawl-mode: folder): Starts at a designated root path (e.g., /Company Home/Sites) or folder UUID and recursively discovers children. Highly configurable with excluded-folder-paths (e.g., ignoring /System or /Sites/trash) to prevent scraping irrelevant files.
  2. CMIS-SQL Relational Queries (crawl-mode: query): Ingests documents dynamically using standard CMIS-SQL statements (e.g. SELECT * FROM cmis:document WHERE cmis:creationDate >= '2026-01-01T00:00:00.000Z'). Allows precision ingestion filtered by metadata criteria, mime types, or custom types.
  3. Incremental Change Log Delta Sync (change-log-enabled: true): Queries the repository's audit log (getContentChanges) using opaque change tokens, efficiently emitting incremental updates and automated OIS Tombstones for deleted objects.

4. Secondary Types, Aspects & Versioning Policies

Enterprise ECM repositories enrich core documents with secondary types (commonly known as aspects in Alfresco and Documentum). The CMIS connector automatically extracts:

Additionally, the connector provides configurable Versioning Policies:

5. Zero-Trust Security: Mapping CMIS ACEs into OIS Permissions

Enterprise AI assistants cannot deliver value if they compromise document confidentiality. If a user queries the AI for confidential acquisition memos, the LLM must only cite documents the user is explicitly permitted to read in the underlying ECM.

The CmisSecurityMapper inspects repository Access Control Entries (ACEs) and translates them into Open Ingestion Standard (OIS) SecurityConfig rules:

{
  "id": "cmis://-default-/documents/e4610814-d865-4907-a108-14d86559076d",
  "action": "UPSERT",
  "metadata": {
    "cmis.objectId": "e4610814-d865-4907-a108-14d86559076d",
    "cmis.name": "enterprise_architecture_2026.pdf",
    "cmis.versionLabel": "2.0",
    "cmis.secondaryObjectTypeIds": ["P:cm:titled", "P:cm:author"],
    "cm:title": "Enterprise Cloud Architecture Roadmap"
  },
  "security": {
    "isExact": true,
    "permissions": [
      { "principal": "GROUP_CLOUD_ARCHITECTS", "identityType": "group", "access": "read" },
      { "principal": "admin", "identityType": "user", "access": "admin" },
      { "principal": "everyone", "identityType": "public", "access": "read" }
    ]
  }
}

These permissions are propagated downstream to vector search stores (PGVector, Milvus, Qdrant, Solr 10, Luxir) and enforced at query time via OpenCrawling's Secure Model Context Protocol (MCP) Server.

6. Claim Check Pattern for Binary Content Streams

Enterprise repositories frequently store large PDF documents, technical schematics, and multimedia assets ranging from tens to hundreds of megabytes. Pushing large binary payloads directly into message brokers like Apache Kafka causes broker latency and disk saturation.

The CMIS connector seamlessly integrates with OpenCrawling's Claim Check Pattern:

  1. The connector streams the binary content stream (cmisselector=content) into the configured Claim Check store (local shared volume, Apache Ozone, or AWS S3).
  2. It generates a lightweight storage URI reference (e.g., file:///data/claims/doc-uuid.pdf or s3://claims/doc-uuid.pdf).
  3. Only the metadata and Claim Check URI reference are dispatched across Kafka, keeping topic traffic ultra-fast and lightweight.
  4. Downstream oc-ingestion-consumer workers fetch the binary directly from storage and extract text using Apache Tika.

7. Configuration & Deployment

Configuring the connector is straightforward via application.yml or environment variables:

Property Key Environment Variable Default Description
spring.opencrawling.connector.cmis.endpoint-url CMIS_ENDPOINT_URL http://localhost:8080/alfresco/.../browser CMIS Browser Binding endpoint URL
spring.opencrawling.connector.cmis.repository-id CMIS_REPOSITORY_ID "" (auto-discovered) Target repository identifier
spring.opencrawling.connector.cmis.username CMIS_USERNAME admin Repository authentication username
spring.opencrawling.connector.cmis.password CMIS_PASSWORD admin Repository authentication password
spring.opencrawling.connector.cmis.crawl-mode CMIS_CRAWL_MODE folder Crawl strategy: folder or query
spring.opencrawling.connector.cmis.root-folder-path CMIS_ROOT_FOLDER_PATH / Starting repository folder path
spring.opencrawling.connector.cmis.versions-mode CMIS_VERSIONS_MODE latest_major Version filter: latest_major, latest, all
spring.opencrawling.connector.cmis.include-acls CMIS_INCLUDE_ACLS true Map CMIS ACEs into OIS security rules

Docker Compose Decoupled Example

services:
  oc-crawler:
    image: opencrawling/oc-runtime:latest
    environment:
      CONNECTOR_TYPE: cmis
      SCAN_PATH: /Company Home
      CMIS_ENDPOINT_URL: http://alfresco:8080/alfresco/api/-default-/public/cmis/versions/1.1/browser
      CMIS_REPOSITORY_ID: -default-
      CMIS_USERNAME: admin
      CMIS_PASSWORD: admin
      CMIS_CRAWL_MODE: folder
      CMIS_INCLUDE_SUBFOLDERS: "true"
      CMIS_VERSIONS_MODE: latest_major
      CMIS_INCLUDE_ACLS: "true"
      CMIS_INCLUDE_CONTENT_STREAM: "true"
      SPRING_OPENCRAWLING_CLAIM_CHECK_STORE: local
      SPRING_OPENCRAWLING_CLAIM_CHECK_LOCAL_DIR: /data/claims
      KAFKA_BOOTSTRAP_SERVERS: kafka:9092
    volumes:
      - shared-data:/data/claims

8. Automated Verification & Testing

The connector is accompanied by automated test suites validating both standalone and distributed decoupled execution against mock and real CMIS services:

✅

1-Command Decoupled Integration Test: Run ./scripts/test-cmis-decoupled.sh or ./run-integration-tests.sh scripts/test-cmis-decoupled.sh to spin up the complete end-to-end containerized environment, execute a live crawl, verify Kafka event streams, and validate semantic vector queries in PostgreSQL PGVector.

Ready to Connect Your Enterprise Content to AI?

Test the brand new CMIS source in our live interactive simulator, explore the full documentation on our Wiki, or clone the repository on GitHub.