1. The Dilemma: AI Vector Ingestion vs. Large-Scale Content Migration
Until today, OpenCrawling’s architecture was laser-focused on Retrieval-Augmented Generation (RAG). In a standard RAG pipeline, crawled documents undergo extensive processing:
- Parsing documents through Apache Tika 4.1.0 inside process-isolated worker JVMs.
- Translating schemas into Markdown with the Auto-Narrativization Copilot.
- Splitting narratives and text into token chunks via
TokenTextSplitter. - Dispatched batches to embedding models (Ollama, OpenAI, HuggingFace) across decoupled
oc-embedding-serviceinstances. - Indexing dense vector embeddings and metadata into vector engines (pgvector, Solr 10, Qdrant, Vespa, Luxir).
While this pipeline is essential for semantic search and generative AI copilots, enterprise IT teams faced a major hurdle when undertaking bulk content migration, platform decommissioning, or cloud data lakehouse archiving.
In migration scenarios, transforming content into text chunks and vectors destroys the original format. Organizations need the pristine, bit-for-bit binary file (CAD designs, medical imaging DICOMs, legal PDFs with digital certificates, TIFF scans), coupled with complete metadata lineage, cryptographic SHA-256 integrity checksums, and source Access Control Lists (ACLs). Running RAG processing on such workloads introduces massive compute bottlenecks without adding migration value.
2. Dual Pipeline Modes: RAG vs. Migration Mode
To solve this impedance mismatch, OpenCrawling introduces the first-class PipelineMode abstraction (RAG vs. MIGRATION):
| Feature | RAG Mode (rag, default) |
Migration Mode (migration) |
|---|---|---|
| Primary Objective | Semantic search & LLM generative context | Bit-for-bit content migration & data liberation |
| Tika Text Extraction | Executed (PipesForkParser isolation) | Bypassed completely (0 CPU cycles) |
| Narrativization Copilot | Executed (Mustache templates) | Bypassed completely |
| Chunking & Splitting | Executed (TokenTextSplitter) | Bypassed completely |
| Vector Embeddings | Generated by oc-embedding-service |
Bypassed completely |
| Destination Payload | Dense vector embeddings + chunk text | Original binary + OIS JSON Sidecar (.ois.json) |
| Target Sinks | pgvector, Solr 10, Qdrant, Vespa, Luxir, SeaTunnel | Apache Ozone (Volumes & Buckets) |
How the Mode is Configured
OpenCrawling enables flexible mode resolution across all interfaces:
- Global System Default: Set
opencrawling.pipeline.mode=migrationin yourapplication.ymlor environment variableOPENCRAWLING_PIPELINE_MODE=migration. - Per-Job Granularity: Override on a per-job basis. A single OpenCrawling cluster can run RAG indexing jobs for knowledge base discovery concurrently with high-speed migration jobs for archival.
- Admin UI Visual Badges: The Job Creation modal features a dedicated Pipeline Mode radio toggle (RAG vs. Migration). Selecting an Apache Ozone output automatically switches the job to Migration mode, adorned with an amber Archive indicator.
- OpenCrawling CLI: Launch jobs directly via
oc job start --name "Legal-Archive" --connector cmis --path "/Company Home/Legal" --mode migration. - Java Client SDK: Strongly typed builder support via
JobRequest.builder().pipelineMode("migration").build().
Strict Fail-Fast Protection: To prevent accidental misconfigurations, attempting to run the Apache Ozone Output Connector in RAG mode fails fast with an immediate IllegalStateException, and REST API submissions return HTTP 409 Conflict ("Rejected: Ozone output requires migration mode").
3. Bit-for-Bit Parity & Open Ingestion Standard (OIS) Metadata Sidecars
When a migration job executes, documents are streamed directly through the Claim Check Pattern into target Apache Ozone volumes and buckets.
For every migrated document, OpenCrawling writes two paired objects:
- The Content Binary: Placed at
<key>(e.g.contracts/2026/master_agreement.pdf), preserving the exact file size, byte sequence, and MIME type. - The OIS Metadata Sidecar: Placed at
<key>.ois.json, encapsulating the Open Ingestion Standard zero-trust envelope.
Here is an example of an emitted .ois.json metadata sidecar:
{
"documentId": "cmis-doc-90421",
"contentRef": "ofs://opencrawling/migration/contracts/2026/master_agreement.pdf",
"checksumSha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
"sourceSystem": "alfresco-content-services",
"mimeType": "application/pdf",
"size": 14829104,
"lastModified": "2026-10-06T14:30:00Z",
"metadata": {
"cmis:name": "master_agreement.pdf",
"cmis:versionLabel": "2.1",
"contract:status": "EXECUTED",
"contract:counterparty": "Acme Global Corp"
},
"security": {
"inheritanceEnabled": true,
"permissions": [
{ "type": "USER", "authority": "alice", "accessType": "READ" },
{ "type": "GROUP", "authority": "legal-counsel", "accessType": "READ_WRITE" }
],
"accessTokens": ["user:alice", "group:legal-counsel"]
}
}
By writing sidecars as companion JSON objects in the same storage bucket, target systems, compliance auditors, and downstream search engines can inspect or ingest document lineage and security rules without inspecting or tampering with the raw binary.
4. Dual Transport Strategies: Native RPC vs. S3 Gateway
oc-ozone-output-connector provides dual client implementations to seamlessly fit any enterprise networking topology:
A. Ozone Native RPC Client (NATIVE)
Communicates directly with the OzoneManager (OM) and storage Datanodes via Apache Hadoop/Ozone RPC protocols (port 9862) using OzoneNativeStorageClient and ofs:// file system schemes.
- Maximum Performance: Eliminates HTTP gateway serialization overhead, providing direct block stream writes to Datanodes.
- Connection Pooling & Caching: Reuses OM connection sessions across document write operations.
- Throughput Advantage: Consistently ~30% faster than HTTP gateway proxies.
B. Ozone S3 Gateway Client (S3G)
Communicates with Apache Ozone via its built-in, AWS S3-compatible HTTP REST Gateway (port 9878) using modern AWS SDK v2.
- Simplified Networking: Only requires an HTTP endpoint exposed to client workers, making it ideal for multi-VPC, Kubernetes egress, or restricted DMZ architectures.
- Parallel Multipart Uploads (
S3MultipartUploader): For multi-gigabyte files exceedingmultipart-threshold: 256MB, the connector automatically forks concurrent part uploads (multipart-part-size: 16MB,multipart-concurrency: 4) to maximize network bandwidth.
5. Parallel Decoupled Kafka Architecture
Enterprise migration pipelines must scale horizontally without encountering deadlocks or queue starvation. OpenCrawling achieves this via decoupled Kafka consumers:
Dedicated Migration Consumer Group: The decoupled writer runs under OzoneMigrationWriterConsumer with a dedicated consumer group (opencrawling-ozone-migration-group). Because this group is separate from the RAG ingestion consumer group, migration traffic never competes for partition leases with RAG indexing workloads.
Key architectural performance highlights include:
- Document-Keyed Partition Ordering: Ingestion messages in Kafka are keyed by
documentId. All lifecycle events for a specific document (e.g. initialUPSERTfollowed by a subsequentDELETEtombstone) land on the same partition and are guaranteed to execute in chronological order. - Virtual Thread Crawler Lanes: The standalone crawler service (
oc-crawler) parallelizes source scanning across concurrent lanes (spring.opencrawling.crawler.concurrency: 4). Connectors supporting lazy streaming (such asFileSystemRepositoryConnectorviaLazyFileInputStream) buffer and upload claim checks concurrently. - Configurable Worker Concurrency: Scale writer parallelism via
spring.opencrawling.output.ozone.consumer-concurrency: 3to process topic partitions in parallel.
6. Document Lifecycles: Deletion Tombstones & Archival
Migrations are rarely one-time static operations; long-running migration waves must handle source deletions gracefully. Supporting the Open Ingestion Standard (OIS) document lifecycle specification, the connector offers configurable tombstone handling via tombstone-action:
DELETE_KEY: Deletes both the target object binary and its.ois.jsonsidecar from Apache Ozone when a source deletion event is processed.ARCHIVE_TOMBSTONE: Deletes the binary object to liberate storage capacity while preserving or appending a tombstone record in the.ois.jsonmetadata sidecar to maintain compliance audit trails.
7. Performance Benchmarks
We measured end-to-end migration throughput using a standardized 400-document dataset (131 MB of mixed PDFs, Word docs, and office files ranging from 4 KB to 1 MB) against Apache Ozone 2.2.0:
| Transport Client | Concurrency (Threads / Partitions) | Throughput (Docs/sec) | Data Transfer Rate | Relative Speedup |
|---|---|---|---|---|
S3 Gateway (S3G) |
1 thread | 19.1 docs/s | 6.3 MB/s | 1.0× (baseline) |
S3 Gateway (S3G) |
3 threads | 33.3 docs/s | 11.0 MB/s | 1.75× |
S3 Gateway (S3G) |
6 threads | 38.4 docs/s | 12.6 MB/s | 2.00× |
Ozone Native RPC (NATIVE) |
1 thread | 25.2 docs/s | 8.3 MB/s | 1.32× |
Ozone Native RPC (NATIVE) |
3 threads | 42.2 docs/s | 13.9 MB/s | 2.21× |
Ozone Native RPC (NATIVE) |
6 threads | 51.1 docs/s | 16.8 MB/s | 2.68× |
Across all benchmarks, the Native RPC client is consistently ~30% faster than the S3 Gateway by bypassing HTTP proxying, while multi-threaded Kafka partition consumption doubles migration velocity.
8. Quick Configuration Reference
To configure Apache Ozone Migration Mode in your Spring Boot environment, add the following to application.yml:
spring:
opencrawling:
output:
type: ozone
ozone:
client-type: NATIVE # NATIVE (ofs/RPC) or S3G (S3 Gateway HTTP)
volume: opencrawling
bucket: migration
om-host: localhost
om-port: 9862
s3-endpoint: http://localhost:9878
access-key: any
secret-key: any
auto-create-bucket: true
key-strategy: HIERARCHICAL # HIERARCHICAL or FLAT
sidecar-suffix: .ois.json
tombstone-action: DELETE_KEY # DELETE_KEY or ARCHIVE_TOMBSTONE
consumer-group: opencrawling-ozone-migration-group
consumer-concurrency: 3
opencrawling:
pipeline:
mode: migration # Activates Migration Mode (bypasses RAG)
| Property Key | Default | Description |
|---|---|---|
opencrawling.pipeline.mode |
rag |
Global pipeline mode: rag or migration |
spring.opencrawling.output.type |
"" |
Set to ozone to activate the Ozone output connector |
spring.opencrawling.output.ozone.client-type |
NATIVE |
Transport strategy: NATIVE (RPC ofs) or S3G (HTTP S3 Gateway) |
spring.opencrawling.output.ozone.volume |
opencrawling |
Target Apache Ozone volume name |
spring.opencrawling.output.ozone.bucket |
migration |
Target Apache Ozone bucket name |
spring.opencrawling.output.ozone.om-host |
localhost |
OzoneManager host for native RPC transport |
spring.opencrawling.output.ozone.om-port |
9862 |
OzoneManager RPC port |
spring.opencrawling.output.ozone.s3-endpoint |
http://localhost:9878 |
Ozone S3 Gateway HTTP endpoint |
spring.opencrawling.output.ozone.multipart-threshold |
256MB |
Object size threshold to trigger parallel S3 multipart upload |
spring.opencrawling.output.ozone.consumer-concurrency |
3 |
Number of parallel decoupled Kafka consumer writer threads |
spring.opencrawling.output.ozone.tombstone-action |
DELETE_KEY |
Lifecycle deletion behavior: DELETE_KEY or ARCHIVE_TOMBSTONE |
9. Automated Verification & Testing
The Migration Mode pipeline is verified via automated test suites in the repository:
Single-Node Integration Test: Run ./scripts/test-ozone-migration.sh to test Ozone bucket creation, binary uploading, and OIS sidecar validation against an embedded runtime.
Decoupled Pipeline Test: Run ./scripts/test-ozone-migration-decoupled.sh to spin up the complete containerized stack: Apache Ozone cluster (SCM, OM, Datanode, S3G) • Kafka broker • oc-crawler (concurrent lanes) • OzoneMigrationWriterConsumer • offline schema validation via CLI.
Ready to Accelerate Your Enterprise Content Migrations?
Explore Migration Mode in our live interactive simulator, read the comprehensive guide on our Wiki, or clone the repository on GitHub.