Connector Release October 5, 2026 • 8 min read

Universal Relational Ingestion & Tabular RAG: Announcing the JDBC Repository Connector

While enterprise content management (ECM) platforms and document stores hold millions of corporate files, the overwhelming majority of day-to-day enterprise intelligence resides directly in Relational Database Management Systems (RDBMS): ERP invoices, CRM client notes, support case tickets, catalog entries, and binary document attachments stored inside database columns. Today, we are proud to release oc-jdbc-repository-connector, unlocking universal relational database crawling, schema-aware Tabular RAG narrativization, binary BLOB streaming, and column-based Zero-Trust Access Control Lists (ACLs) for OpenCrawling! 🎉

JDBC Relational Ingestion & Tabular RAG Pipeline Architecture
RDBMS (Postgres / MySQL / Oracle / DB2)
→
HikariCP & Streaming Cursor (Java 25)
→
StructuredTaskScope Batching
→
Dual Path: Tabular RAG / BLOB Claim Check
→
Kafka Broker
→
Vector Stores & Secure MCP Server

1. The Impedance Mismatch: Relational Data vs. Vector RAG

Building Retrieval-Augmented Generation (RAG) systems over relational databases has historically presented an awkward engineering trade-off. Naive approaches often flatten tables into dense comma-separated values (CSVs) or dump raw JSON structures into text splitters.

This approach causes two critical issues:

The JDBC Repository Connector bridges this divide. It introduces native Dual-Path Ingestion: transforming relational records into rich natural-language narratives for Tabular RAG while routing binary attachments (PDFs, images, office files) directly through Apache Tika via decoupled Claim-Check streaming.

2. Universal SQL Compatibility Across Major Database Engines

Built on standard JDBC 4.2+ specifications and enterprise-grade HikariCP connection pooling, oc-jdbc-repository-connector runs out-of-the-box against all major enterprise database engines:

💡

Engine-Specific Streaming Optimizations: To guarantee constant $O(1)$ client heap utilization when scanning multi-million-row tables, the connector automatically adjusts JDBC driver behavior. In PostgreSQL, it disables auto-commit (connection.setAutoCommit(false)) to activate server-side cursor streaming; in MySQL, it sets streaming fetch mode (statement.setFetchSize(Integer.MIN_VALUE)) to prevent driver buffering.

3. Dual Ingestion Strategies: Table Mode vs. Custom SQL Query Mode

The connector offers two operational strategies configurable via the Admin UI, Spring Boot YAML, or dynamic SCAN_PATH parameters:

A. Table / View Ingestion Mode

Specify a target table or view (e.g. support_tickets or crm.customers). The connector inspects table metadata via DatabaseMetaData.getColumns(), identifies primary keys, and projects all columns into Open Ingestion Standard (OIS) attributes.

B. Custom SQL Query Mode

For complex domain models involving normalized schemas, supply an arbitrary SQL query joining multiple tables:

SELECT t.id AS ticket_id, t.title, t.description, t.status, t.created_at, t.updated_at,
       c.name AS customer_name, c.tier AS customer_tier,
       t.owner_username, t.department_id, t.is_deleted, t.tenant_id
FROM support_tickets t
LEFT JOIN customers c ON t.customer_id = c.id
WHERE (:lastCrawledTime IS NULL OR t.updated_at >= :lastCrawledTime)

C. Multi-Table Concurrent Scanning

Need to ingest an entire database catalog at once? The connector accepts a comma-separated list of tables (e.g. customers,orders,invoices). It forks parallel scanning tasks across Java 25 StructuredTaskScope subtasks, maximizing I/O throughput across database connections.

4. Tabular RAG & Auto-Narrativization Copilot

OpenCrawling's RepositoryConnector.getSchema(basePath) SPI inspects SQL columns, data types, and remarks at runtime. This feeds directly into the Admin UI and the Spring AI TemplateGenerationCopilot, which automatically synthesizes natural-language Mustache templates.

For example, given a support ticket row, the narrativizer produces:

# Support Ticket {{id}}: {{title}}
Customer: {{customer_name}} (Tier: {{customer_tier}})
Status: {{status}} | Assigned to: {{owner_username}}
Created: {{created_at}} | Last Updated: {{updated_at}}

## Description
{{description}}

When narrativization is disabled, the connector generates a clean, deterministic Markdown key-value fallback representation, eliminating the flat CSV anti-pattern and maximizing vector retrieval relevance.

5. Binary BLOB / CLOB Streaming with Magic Byte Sniffing

Many enterprise databases store binary files directly in database columns (such as BLOB, BYTEA, RAW, or LONGVARBINARY). Loading these binaries into memory during database cursor scans is a notorious cause of JVM OutOfMemoryErrors.

The JDBC connector solves this via OpenCrawling's decoupled Claim Check Pattern:

6. Incremental Delta Crawling & Soft Deletes

Enterprise databases change continuously. The JDBC connector provides full change data capture (CDC) mechanisms:

7. Zero-Trust Security & Column-Based ACL Mapping

Security is a first-class citizen in OpenCrawling. Using JdbcSecurityMapper, database security columns are mapped directly into standard OIS SecurityConfig and PermissionRule objects:

Downstream, OpenCrawling's Secure Model Context Protocol (MCP) Server enforces strict principal filtering against vector store indices, guaranteeing that users and LLM agents only retrieve information they have explicit row-level authority to view.

8. Quick Configuration Reference

Add the following configuration to your application.yml or environment variables:

Property Key Environment Variable Default Description
spring.opencrawling.connector.jdbc.url JDBC_URL jdbc:h2:mem:... JDBC connection string (PostgreSQL, MySQL, Oracle, etc.)
spring.opencrawling.connector.jdbc.driver-class-name JDBC_DRIVER_CLASS_NAME "" (auto-detected) Driver class name (e.g. org.postgresql.Driver)
spring.opencrawling.connector.jdbc.username JDBC_USERNAME sa Database authentication username
spring.opencrawling.connector.jdbc.password JDBC_PASSWORD "" Database authentication password
spring.opencrawling.connector.jdbc.mode JDBC_MODE table Ingestion mode: table or query
spring.opencrawling.connector.jdbc.table-name JDBC_TABLE_NAME "" Table or view name to ingest
spring.opencrawling.connector.jdbc.primary-key-columns JDBC_PRIMARY_KEY_COLUMNS id Comma-separated column name(s) for document primary key
spring.opencrawling.connector.jdbc.title-column JDBC_TITLE_COLUMN title Column mapped to document title
spring.opencrawling.connector.jdbc.blob-column-name JDBC_BLOB_COLUMN_NAME "" Binary BLOB column name (auto-detected if blank)
spring.opencrawling.connector.jdbc.security-enabled JDBC_SECURITY_ENABLED false Enable row security and column-to-ACL mapping
spring.opencrawling.connector.jdbc.user-columns JDBC_USER_COLUMNS "" User identity columns (e.g. owner_id)
spring.opencrawling.connector.jdbc.group-columns JDBC_GROUP_COLUMNS "" Group/role columns (e.g. department_id)
spring.opencrawling.connector.jdbc.soft-delete-enabled JDBC_SOFT_DELETE_ENABLED false Enable soft-delete tombstone emission

9. Automated Verification & Testing

The JDBC connector is backed by comprehensive integration testing scripts:

✅

Standalone Test Script: Run ./scripts/test-jdbc-connector.sh to test schema introspection, BLOB extraction, and structured concurrency across in-memory H2 and ephemeral PostgreSQL Testcontainers.

✅

Decoupled Pipeline Test: Run ./scripts/test-jdbc-decoupled.sh to execute the full end-to-end containerized pipeline: PostgreSQL source database → oc-crawler (JDBC) → Kafka broker → Ingestion Consumer (Tika & chunks) → Embedding Consumer (Ollama) → Writer Consumer (pgvector) → Secure MCP Server queries.

Ready to Connect Your Relational Data to AI?

Test the brand new JDBC Repository Connector in our interactive simulator, read the comprehensive documentation on our Wiki, or clone the repository on GitHub.