1. The Case for an Open Standard in Enterprise AI
Building Retrieval-Augmented Generation (RAG) applications in Fortune 500 and public-sector environments rarely involves a single, clean content repository. Large enterprises often maintain historical contracts in IBM FileNet, regulated engineering designs in OpenText Documentum, collaborative project spaces in Alfresco Content Services, and departmental archives in Nuxeo.
Creating bespoke, vendor-specific crawlers for every proprietary ECM API leads to technical debt, mismatched metadata schemas, and fragile security boundaries.
The Content Management Interoperability Services (CMIS) specificationâstandardized by OASISâwas created to solve exactly this challenge. By adopting CMIS as a universal abstraction layer, OpenCrawling enables organizations to connect once and ingest from any compliant ECM system, preserving folder hierarchies, custom aspects, document versions, and fine-grained permissions without vendor lock-in.
2. Pure Java 25 & Zero Legacy Dependencies
Historically, Java integrations with CMIS relied heavily on the Apache Chemistry OpenCMIS library. While Apache Chemistry pioneered CMIS adoption, its client libraries carry dozens of legacy transitive dependencies, heavy XML binding frameworks (JAXB/JAX-WS), and synchronous I/O architectures ill-suited for modern reactive microservices and containerized runtimes.
OpenCrawling's oc-cmis-repository-connector is built with a fundamentally modern approach:
-
Native Java 25
HttpClient: Implements the CMIS 1.1 Browser Binding directly over HTTP/JSON using Java 25's native HTTP client, reducing JAR size, eliminating security vulnerabilities, and speeding container boot times to sub-second levels. -
Java 25 Virtual Threads (
StructuredTaskScope): Recursive folder traversal, pagination batches, and metadata enrichment execute concurrently across lightweight virtual threads with zero risk of thread pool exhaustion. -
Succinct & Verbose JSON Parsing: Efficiently consumes both CMIS 1.1 concise (
succinctProperties) and traditional verbose property formats using Jackson Streaming.
Zero Chemistry Footprint: By eliminating legacy Chemistry JARs, oc-cmis-repository-connector runs cleanly on modular Java 25 runtimes, consumes up to 70% less heap memory during massive scans, and supports instant container scaling.
3. Multi-Mode Crawling: Hierarchies, Queries & Change Logs
Different enterprise search and compliance use cases demand different ingestion strategies. The CMIS connector supports three distinct crawl modes:
-
Recursive Folder Crawl (
crawl-mode: folder): Starts at a designated root path (e.g.,/Company Home/Sites) or folder UUID and recursively discovers children. Highly configurable withexcluded-folder-paths(e.g., ignoring/Systemor/Sites/trash) to prevent scraping irrelevant files. -
CMIS-SQL Relational Queries (
crawl-mode: query): Ingests documents dynamically using standard CMIS-SQL statements (e.g.SELECT * FROM cmis:document WHERE cmis:creationDate >= '2026-01-01T00:00:00.000Z'). Allows precision ingestion filtered by metadata criteria, mime types, or custom types. -
Incremental Change Log Delta Sync (
change-log-enabled: true): Queries the repository's audit log (getContentChanges) using opaque change tokens, efficiently emitting incremental updates and automated OIS Tombstones for deleted objects.
4. Secondary Types, Aspects & Versioning Policies
Enterprise ECM repositories enrich core documents with secondary types (commonly known as aspects in Alfresco and Documentum). The CMIS connector automatically extracts:
- Secondary Types List: Ingests
cmis:secondaryObjectTypeIds(e.g.P:cm:titled,P:cm:author,P:custom:confidentiality). - Dynamic Custom Metadata: All custom aspect properties attached to documents are mapped into the OIS document payload, making them fully searchable and filterable in downstream vector stores.
Additionally, the connector provides configurable Versioning Policies:
latest_major: Indexes only official approved revisions (cmis:isLatestMajorVersion = true), ideal for corporate policies and product documentation.latest: Indexes the newest working draft (cmis:isLatestVersion = true).all: Indexes all revisions in the version series for compliance audits.
5. Zero-Trust Security: Mapping CMIS ACEs into OIS Permissions
Enterprise AI assistants cannot deliver value if they compromise document confidentiality. If a user queries the AI for confidential acquisition memos, the LLM must only cite documents the user is explicitly permitted to read in the underlying ECM.
The CmisSecurityMapper inspects repository Access Control Entries (ACEs) and translates them into Open Ingestion Standard (OIS) SecurityConfig rules:
{
"id": "cmis://-default-/documents/e4610814-d865-4907-a108-14d86559076d",
"action": "UPSERT",
"metadata": {
"cmis.objectId": "e4610814-d865-4907-a108-14d86559076d",
"cmis.name": "enterprise_architecture_2026.pdf",
"cmis.versionLabel": "2.0",
"cmis.secondaryObjectTypeIds": ["P:cm:titled", "P:cm:author"],
"cm:title": "Enterprise Cloud Architecture Roadmap"
},
"security": {
"isExact": true,
"permissions": [
{ "principal": "GROUP_CLOUD_ARCHITECTS", "identityType": "group", "access": "read" },
{ "principal": "admin", "identityType": "user", "access": "admin" },
{ "principal": "everyone", "identityType": "public", "access": "read" }
]
}
}
These permissions are propagated downstream to vector search stores (PGVector, Milvus, Qdrant, Solr 10, Luxir) and enforced at query time via OpenCrawling's Secure Model Context Protocol (MCP) Server.
6. Claim Check Pattern for Binary Content Streams
Enterprise repositories frequently store large PDF documents, technical schematics, and multimedia assets ranging from tens to hundreds of megabytes. Pushing large binary payloads directly into message brokers like Apache Kafka causes broker latency and disk saturation.
The CMIS connector seamlessly integrates with OpenCrawling's Claim Check Pattern:
- The connector streams the binary content stream (
cmisselector=content) into the configured Claim Check store (local shared volume, Apache Ozone, or AWS S3). - It generates a lightweight storage URI reference (e.g.,
file:///data/claims/doc-uuid.pdfors3://claims/doc-uuid.pdf). - Only the metadata and Claim Check URI reference are dispatched across Kafka, keeping topic traffic ultra-fast and lightweight.
- Downstream
oc-ingestion-consumerworkers fetch the binary directly from storage and extract text using Apache Tika.
7. Configuration & Deployment
Configuring the connector is straightforward via application.yml or environment variables:
| Property Key | Environment Variable | Default | Description |
|---|---|---|---|
spring.opencrawling.connector.cmis.endpoint-url |
CMIS_ENDPOINT_URL |
http://localhost:8080/alfresco/.../browser |
CMIS Browser Binding endpoint URL |
spring.opencrawling.connector.cmis.repository-id |
CMIS_REPOSITORY_ID |
"" (auto-discovered) |
Target repository identifier |
spring.opencrawling.connector.cmis.username |
CMIS_USERNAME |
admin |
Repository authentication username |
spring.opencrawling.connector.cmis.password |
CMIS_PASSWORD |
admin |
Repository authentication password |
spring.opencrawling.connector.cmis.crawl-mode |
CMIS_CRAWL_MODE |
folder |
Crawl strategy: folder or query |
spring.opencrawling.connector.cmis.root-folder-path |
CMIS_ROOT_FOLDER_PATH |
/ |
Starting repository folder path |
spring.opencrawling.connector.cmis.versions-mode |
CMIS_VERSIONS_MODE |
latest_major |
Version filter: latest_major, latest, all |
spring.opencrawling.connector.cmis.include-acls |
CMIS_INCLUDE_ACLS |
true |
Map CMIS ACEs into OIS security rules |
Docker Compose Decoupled Example
services:
oc-crawler:
image: opencrawling/oc-runtime:latest
environment:
CONNECTOR_TYPE: cmis
SCAN_PATH: /Company Home
CMIS_ENDPOINT_URL: http://alfresco:8080/alfresco/api/-default-/public/cmis/versions/1.1/browser
CMIS_REPOSITORY_ID: -default-
CMIS_USERNAME: admin
CMIS_PASSWORD: admin
CMIS_CRAWL_MODE: folder
CMIS_INCLUDE_SUBFOLDERS: "true"
CMIS_VERSIONS_MODE: latest_major
CMIS_INCLUDE_ACLS: "true"
CMIS_INCLUDE_CONTENT_STREAM: "true"
SPRING_OPENCRAWLING_CLAIM_CHECK_STORE: local
SPRING_OPENCRAWLING_CLAIM_CHECK_LOCAL_DIR: /data/claims
KAFKA_BOOTSTRAP_SERVERS: kafka:9092
volumes:
- shared-data:/data/claims
8. Automated Verification & Testing
The connector is accompanied by automated test suites validating both standalone and distributed decoupled execution against mock and real CMIS services:
1-Command Decoupled Integration Test: Run ./scripts/test-cmis-decoupled.sh or ./run-integration-tests.sh scripts/test-cmis-decoupled.sh to spin up the complete end-to-end containerized environment, execute a live crawl, verify Kafka event streams, and validate semantic vector queries in PostgreSQL PGVector.
Ready to Connect Your Enterprise Content to AI?
Test the brand new CMIS source in our live interactive simulator, explore the full documentation on our Wiki, or clone the repository on GitHub.