Enterprise Migration October 6, 2026 • 8 min read

Bit-for-Bit Enterprise Data Liberation: Announcing Migration Mode and the Apache Ozone Output Connector

Enterprise data liberation is fundamentally different from AI ingestion. When moving petabytes of proprietary content out of legacy enterprise content management (ECM) repositories and into modern, cloud-native object storage, organizations need bit-for-bit fidelity, cryptographic verification, and strict Zero-Trust access control list (ACL) preservation. They cannot afford to waste CPU cycles and compute budgets extracting text, generating narrative templates, or calculating vector embeddings. Today, we are thrilled to introduce Migration Mode and the Apache Ozone Output Connector (oc-ozone-output-connector) to OpenCrawling! 🚀

Migration Mode & Apache Ozone Output Pipeline Architecture
Source Repository (Alfresco / CMIS / JDBC / Filesystem)
→
Crawler (Parallel Lanes + Claim-Check)
→
Kafka Topic (opencrawling-documents)
→
Bypass: No Tika • No Embeddings
→
OzoneMigrationWriterConsumer
→
Apache Ozone: Pristine Binary + .ois.json Sidecar

1. The Dilemma: AI Vector Ingestion vs. Large-Scale Content Migration

Until today, OpenCrawling’s architecture was laser-focused on Retrieval-Augmented Generation (RAG). In a standard RAG pipeline, crawled documents undergo extensive processing:

While this pipeline is essential for semantic search and generative AI copilots, enterprise IT teams faced a major hurdle when undertaking bulk content migration, platform decommissioning, or cloud data lakehouse archiving.

In migration scenarios, transforming content into text chunks and vectors destroys the original format. Organizations need the pristine, bit-for-bit binary file (CAD designs, medical imaging DICOMs, legal PDFs with digital certificates, TIFF scans), coupled with complete metadata lineage, cryptographic SHA-256 integrity checksums, and source Access Control Lists (ACLs). Running RAG processing on such workloads introduces massive compute bottlenecks without adding migration value.

2. Dual Pipeline Modes: RAG vs. Migration Mode

To solve this impedance mismatch, OpenCrawling introduces the first-class PipelineMode abstraction (RAG vs. MIGRATION):

Feature RAG Mode (rag, default) Migration Mode (migration)
Primary Objective Semantic search & LLM generative context Bit-for-bit content migration & data liberation
Tika Text Extraction Executed (PipesForkParser isolation) Bypassed completely (0 CPU cycles)
Narrativization Copilot Executed (Mustache templates) Bypassed completely
Chunking & Splitting Executed (TokenTextSplitter) Bypassed completely
Vector Embeddings Generated by oc-embedding-service Bypassed completely
Destination Payload Dense vector embeddings + chunk text Original binary + OIS JSON Sidecar (.ois.json)
Target Sinks pgvector, Solr 10, Qdrant, Vespa, Luxir, SeaTunnel Apache Ozone (Volumes & Buckets)

How the Mode is Configured

OpenCrawling enables flexible mode resolution across all interfaces:

🛡

Strict Fail-Fast Protection: To prevent accidental misconfigurations, attempting to run the Apache Ozone Output Connector in RAG mode fails fast with an immediate IllegalStateException, and REST API submissions return HTTP 409 Conflict ("Rejected: Ozone output requires migration mode").

3. Bit-for-Bit Parity & Open Ingestion Standard (OIS) Metadata Sidecars

When a migration job executes, documents are streamed directly through the Claim Check Pattern into target Apache Ozone volumes and buckets.

For every migrated document, OpenCrawling writes two paired objects:

  1. The Content Binary: Placed at <key> (e.g. contracts/2026/master_agreement.pdf), preserving the exact file size, byte sequence, and MIME type.
  2. The OIS Metadata Sidecar: Placed at <key>.ois.json, encapsulating the Open Ingestion Standard zero-trust envelope.

Here is an example of an emitted .ois.json metadata sidecar:

{
  "documentId": "cmis-doc-90421",
  "contentRef": "ofs://opencrawling/migration/contracts/2026/master_agreement.pdf",
  "checksumSha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
  "sourceSystem": "alfresco-content-services",
  "mimeType": "application/pdf",
  "size": 14829104,
  "lastModified": "2026-10-06T14:30:00Z",
  "metadata": {
    "cmis:name": "master_agreement.pdf",
    "cmis:versionLabel": "2.1",
    "contract:status": "EXECUTED",
    "contract:counterparty": "Acme Global Corp"
  },
  "security": {
    "inheritanceEnabled": true,
    "permissions": [
      { "type": "USER", "authority": "alice", "accessType": "READ" },
      { "type": "GROUP", "authority": "legal-counsel", "accessType": "READ_WRITE" }
    ],
    "accessTokens": ["user:alice", "group:legal-counsel"]
  }
}

By writing sidecars as companion JSON objects in the same storage bucket, target systems, compliance auditors, and downstream search engines can inspect or ingest document lineage and security rules without inspecting or tampering with the raw binary.

4. Dual Transport Strategies: Native RPC vs. S3 Gateway

oc-ozone-output-connector provides dual client implementations to seamlessly fit any enterprise networking topology:

A. Ozone Native RPC Client (NATIVE)

Communicates directly with the OzoneManager (OM) and storage Datanodes via Apache Hadoop/Ozone RPC protocols (port 9862) using OzoneNativeStorageClient and ofs:// file system schemes.

B. Ozone S3 Gateway Client (S3G)

Communicates with Apache Ozone via its built-in, AWS S3-compatible HTTP REST Gateway (port 9878) using modern AWS SDK v2.

5. Parallel Decoupled Kafka Architecture

Enterprise migration pipelines must scale horizontally without encountering deadlocks or queue starvation. OpenCrawling achieves this via decoupled Kafka consumers:

⚡

Dedicated Migration Consumer Group: The decoupled writer runs under OzoneMigrationWriterConsumer with a dedicated consumer group (opencrawling-ozone-migration-group). Because this group is separate from the RAG ingestion consumer group, migration traffic never competes for partition leases with RAG indexing workloads.

Key architectural performance highlights include:

6. Document Lifecycles: Deletion Tombstones & Archival

Migrations are rarely one-time static operations; long-running migration waves must handle source deletions gracefully. Supporting the Open Ingestion Standard (OIS) document lifecycle specification, the connector offers configurable tombstone handling via tombstone-action:

7. Performance Benchmarks

We measured end-to-end migration throughput using a standardized 400-document dataset (131 MB of mixed PDFs, Word docs, and office files ranging from 4 KB to 1 MB) against Apache Ozone 2.2.0:

Transport Client Concurrency (Threads / Partitions) Throughput (Docs/sec) Data Transfer Rate Relative Speedup
S3 Gateway (S3G) 1 thread 19.1 docs/s 6.3 MB/s 1.0× (baseline)
S3 Gateway (S3G) 3 threads 33.3 docs/s 11.0 MB/s 1.75×
S3 Gateway (S3G) 6 threads 38.4 docs/s 12.6 MB/s 2.00×
Ozone Native RPC (NATIVE) 1 thread 25.2 docs/s 8.3 MB/s 1.32×
Ozone Native RPC (NATIVE) 3 threads 42.2 docs/s 13.9 MB/s 2.21×
Ozone Native RPC (NATIVE) 6 threads 51.1 docs/s 16.8 MB/s 2.68×

Across all benchmarks, the Native RPC client is consistently ~30% faster than the S3 Gateway by bypassing HTTP proxying, while multi-threaded Kafka partition consumption doubles migration velocity.

8. Quick Configuration Reference

To configure Apache Ozone Migration Mode in your Spring Boot environment, add the following to application.yml:

spring:
  opencrawling:
    output:
      type: ozone
      ozone:
        client-type: NATIVE              # NATIVE (ofs/RPC) or S3G (S3 Gateway HTTP)
        volume: opencrawling
        bucket: migration
        om-host: localhost
        om-port: 9862
        s3-endpoint: http://localhost:9878
        access-key: any
        secret-key: any
        auto-create-bucket: true
        key-strategy: HIERARCHICAL       # HIERARCHICAL or FLAT
        sidecar-suffix: .ois.json
        tombstone-action: DELETE_KEY     # DELETE_KEY or ARCHIVE_TOMBSTONE
        consumer-group: opencrawling-ozone-migration-group
        consumer-concurrency: 3

opencrawling:
  pipeline:
    mode: migration                      # Activates Migration Mode (bypasses RAG)
Property Key Default Description
opencrawling.pipeline.mode rag Global pipeline mode: rag or migration
spring.opencrawling.output.type "" Set to ozone to activate the Ozone output connector
spring.opencrawling.output.ozone.client-type NATIVE Transport strategy: NATIVE (RPC ofs) or S3G (HTTP S3 Gateway)
spring.opencrawling.output.ozone.volume opencrawling Target Apache Ozone volume name
spring.opencrawling.output.ozone.bucket migration Target Apache Ozone bucket name
spring.opencrawling.output.ozone.om-host localhost OzoneManager host for native RPC transport
spring.opencrawling.output.ozone.om-port 9862 OzoneManager RPC port
spring.opencrawling.output.ozone.s3-endpoint http://localhost:9878 Ozone S3 Gateway HTTP endpoint
spring.opencrawling.output.ozone.multipart-threshold 256MB Object size threshold to trigger parallel S3 multipart upload
spring.opencrawling.output.ozone.consumer-concurrency 3 Number of parallel decoupled Kafka consumer writer threads
spring.opencrawling.output.ozone.tombstone-action DELETE_KEY Lifecycle deletion behavior: DELETE_KEY or ARCHIVE_TOMBSTONE

9. Automated Verification & Testing

The Migration Mode pipeline is verified via automated test suites in the repository:

✅

Single-Node Integration Test: Run ./scripts/test-ozone-migration.sh to test Ozone bucket creation, binary uploading, and OIS sidecar validation against an embedded runtime.

✅

Decoupled Pipeline Test: Run ./scripts/test-ozone-migration-decoupled.sh to spin up the complete containerized stack: Apache Ozone cluster (SCM, OM, Datanode, S3G) • Kafka broker • oc-crawler (concurrent lanes) • OzoneMigrationWriterConsumer • offline schema validation via CLI.

Ready to Accelerate Your Enterprise Content Migrations?

Explore Migration Mode in our live interactive simulator, read the comprehensive guide on our Wiki, or clone the repository on GitHub.