Enterprise Semantic Lineage & Autonomous Governance

The Enterprise Knowledge Graph
Engineered for Zero Operational Drag

LineagIQ unifies fragmented, multi-cloud data pipelines into an active, queryable semantic knowledge graph. Experience point-in-time schema drift analysis, instant recursive blast-radius attribution, and an ultra-lean serverless architecture that completely eliminates expensive, dedicated graph database clusters.

92%
Reduction in TCO
Replaces 6-figure Neo4j and Neptune clusters with serverless Delta Lake and in-process DuckDB.
< 30s
Root Cause Attribution
Instantly traverses upstream dependencies to pinpoint silent schema breaks, breaking commits, and pipeline failures.
100%
Point-in-Time Auditability
Immutable ACID transactional commit logs delivering full audit readiness for GDPR, BCBS 239, and HIPAA.
0
Heavyweight Daemons
Zero continuous JVM graph servers, zero cluster maintenance overhead, and zero data exfiltration.
LineagIQ Knowledge Graph Visualizer Overview
Live Delta Lake Graph: 37 Entities, 46 Lineage Edges, 3 Historical Time Travel Commits
Historical ACID Time Travel

Eliminate Silent Data Drift With Time-Travel Diffing

Data pipelines fail silently when upstream tables alter column types, rename keys, or drop fields unannounced. LineagIQ captures every structural revision and dependency shift across time through Delta Lake's ACID transactional commit logs.

  • ✓
    Continuous State Reconstruction: Scrub through intuitive timeline controls to recreate and inspect the precise topological state of your data assets at any past timestamp or commit version.
  • ✓
    Automated Topology Diffing (T₁ → T₂): Instantly isolate additions, removals, column mutations, and rerouted lineage dependencies between any two arbitrary points in time.
  • ✓
    Regulatory Audit Readiness: Fulfill stringent financial, governance, and privacy mandates (BCBS 239, SOX, GDPR) with immutable historical proof of upstream provenance and data lineage.
Test Time Travel in Live Demo →
LineagIQ Time Travel Diff Modal
Click to Zoom
Impact & Root-Cause Search

Instant Blast Radius & Root-Cause Attribution

Before deploying schema migrations or refactoring dbt models, LineagIQ computes downstream blast radius in milliseconds—preventing downstream data incidents and broken executive dashboards.

  • ⚡
    Recursive Downstream Traversal: Discover every indirect dependency across data warehouses, feature stores, ML models, reverse-ETL jobs, and BI dashboards across arbitrary graph depths.
  • ⚡
    Upstream Root-Cause Isolation: When an executive KPI or dashboard metric diverges, trace backward through marts, transformations, and ingestion pipelines to uncover the failure source in seconds.
  • ⚡
    Targeted Stakeholder Notification: Map asset ownership, domain boundaries, and downstream business consumers directly onto impacted subgraphs to orchestrate targeted incident response.
Explore Blast Radius in Live Demo →
LineagIQ Downstream Blast Radius Highlight
Click to Zoom
GraphRAG AI Copilot

Autonomous AI Assistant for Architectural Governance

LineagIQ bridges graph intelligence with generative AI through Graph-Augmented Generation (GraphRAG). Ground your LLMs with deep structural lineage context to eliminate hallucinations during architectural audits and incident investigations.

  • 🤖
    Private Local & Cloud LLM Flexibility: Run completely offline with local Ollama instances (Llama 3, Mistral, Qwen) for zero network egress, or connect directly to cloud APIs (OpenAI GPT-4o, Google Gemini).
  • 🤖
    Automated Impact Synthesis: Feed time-travel schema diffs directly into the AI Assistant to synthesize executive incident summaries, risk scores, and remediation playbooks on demand.
  • 🤖
    Standardized Agentic Tooling: Out-of-the-box Python retriever functions for LangChain, AutoGen, and LlamaIndex to query blast radius and historical lineage programmatically.
LineagIQ AI Assistant Drawer
Click to Zoom
Universal Source Ingestion

Inject Any Enterprise Data Source Directly Into the Collection Agent

The LineagIQ Collection Agent operates as a lightweight, stateless utility inside your secure VPC, CI/CD pipeline, or Kubernetes cluster. It ingests schema definitions, warehouse catalogs, audit logs, and runtime orchestrator events—transforming raw operational signals into unified graph models and semantic vector embeddings without ever touching raw customer record data.

⚡

dbt Core & dbt Cloud

Transformation DAGs & Schemas

Directly ingests compilation artifacts (manifest.json and catalog.json). Extracts data models, seeds, snapshots, source declarations, column-level documentation, tags, and tests, compiling the complete transformation DAG into unified DERIVED_FROM dependency edges.

dbt Core dbt Cloud Snowflake BigQuery Databricks
python -m collection_agent.src.cli --dbt-manifest <path> --dbt-catalog <path>
🗄️

SQL Catalogs & Schemas

Warehouse Metadata & Relational DDL

Parses standard INFORMATION_SCHEMA.TABLES, COLUMNS, and foreign key constraint definitions. Extracts relational schemas, column data types, ordinal positions, and physical JOINS_WITH relationships across relational databases and cloud warehouses.

PostgreSQL Snowflake Databricks Unity AWS Redshift Google BigQuery
python -m collection_agent.src.cli --sql-schema <path_to_schema.json>
📜

SQL Query Logs & Audit History

Access Auditing & Runtime Join Mining

Analyzes warehouse audit histories (Snowflake QUERY_HISTORY, BigQuery INFORMATION_SCHEMA.JOBS, Databricks Query History). Mines actual user and service dataset consumption patterns (CONSUMED_BY), query execution frequencies, and implicit multi-table join predicates.

Snowflake History BigQuery Jobs Databricks Logs Athena / CloudTrail
python -m collection_agent.src.cli --query-logs <path_to_logs.json>
🌐

OpenLineage Standard Events

Orchestration & Compute Lineage

Ingests standardized OpenLineage JSON run events emitted by workflow orchestrators and distributed compute engines. Seamlessly maps job run IDs, execution timestamps, input dataset consumption edges (CONSUMED_BY), and output production edges (PRODUCED_BY).

Apache Airflow Apache Spark Dagster Apache Flink dbt-ol
python -m collection_agent.src.cli --openlineage <path_to_events.json>

The "Slim Stack" Revolution: Why No Dedicated Graph DB?

Legacy enterprise data catalogs mandate dedicated Neo4j, TigerGraph, or Amazon Neptune clusters—burdening teams with persistent infrastructure bills, JVM memory management, and specialized graph query skills. LineagIQ re-architects graph intelligence from first principles.

Legacy Data Catalogs

Collibra / Alation / Neo4j Stack

$250k+ / year
  • Heavyweight Infrastructure: 24/7 provisioned JVM graph nodes (Neo4j Enterprise, Amazon Neptune) requiring massive RAM and continuous tuning.
  • Network & Perimeter Risk: Often requires granting third-party SaaS vendors persistent inbound network tunnels or broad warehouse read permissions.
  • Destructive Overwrite Semantics: Traditional graph databases overwrite properties in-place, making historical state recreation cumbersome, slow, or impossible.
  • Proprietary Graph Lock-in: Lineage topologies trapped in niche graph formats and query dialects (Cypher, Gremlin) isolated from standard lakehouse tools.
LineagIQ Slim Stack

Serverless S3 Delta + In-Process DuckDB

< $5k / year
  • Serverless Storage Economics: Topologies persist directly on commodity object storage (Amazon S3, Google Cloud Storage, or local disk) in open Delta Lake Parquet format.
  • Ephemeral Zero-Egress Collection: Collection agents run as lightweight batch jobs inside your VPC or CI/CD pipelines. Operational customer data never leaves your trust boundary.
  • Native ACID Time-Travel Replay: Delta Lake's transactional commit logs capture every lineage revision automatically with zero storage bloat and full historical reproducibility.
  • Vectorized C++ Query Engine: In-process DuckDB evaluates multi-hop recursive graph traversals using standard SQL CTEs in milliseconds directly over columnar Parquet files.
⚡ In-Memory Vectorized Benchmarks DuckDB Snapshot Materialization

Hardware Sizing vs. Graph Scale (Nodes & Edges)

Thanks to LineagIQ's versioned snapshot materialization, historical graph partitions are dynamically cached in DuckDB's in-memory columnar engine. Bit-packed, dictionary-encoded relations achieve sub-10ms recursive CTE query execution on ultra-compact hardware—eliminating expensive multi-node graph database clusters entirely.

Scale Tier Graph Topology (Nodes & Edges) In-Memory Cache Footprint Recommended Hardware Recursive Traversal Latency Legacy JVM Graph Equivalent
Mid-Market 50,000 nodes • 150,000 edges ~25 MB – 45 MB 0.5 vCPU, 1 GB RAM (e.g., ~$10–$20/mo) < 8 ms 16 GB RAM Neo4j (~$300/mo)
Enterprise 250,000 nodes • 1,200,000 edges ~180 MB – 320 MB 1 vCPU, 2 GB RAM (e.g., ~$30/mo) < 18 ms 32 GB RAM Neptune (~$850/mo)
Large Enterprise 1,000,000 nodes • 5,000,000 edges ~650 MB – 1.1 GB 2 vCPU, 4 GB RAM (e.g., ~$50/mo) < 45 ms 64 GB RAM HA Cluster (~$2,200/mo)
Hyperscale Global 5,000,000+ nodes • 25,000,000+ edges ~3.2 GB – 5.8 GB 4 vCPU, 8–16 GB RAM (e.g., ~$90/mo) < 95 ms 128 GB+ Distributed Cluster ($5,000+/mo)

Developer & Architecture Deep Dive

Designed for platform architects, data engineers, and AI developers. LineagIQ cleanly decouples edge metadata extraction, immutable table storage, and vectorized query execution.

Tier 01 // Ephemeral Edge

Stateless Ingestion Agent

Executes as an ephemeral batch task (Docker, Kubernetes CronJob, or GitHub Action) inside customer boundaries. Ingests metadata from files, schemas, query logs, and orchestrators.

dbt manifest.json INFORMATION_SCHEMA QUERY_HISTORY OpenLineage JSON bge-small INT8 ONNX
Tier 02 // Lakehouse Storage

Delta Lake Storage Layer

Standard cloud object store (S3/GCS) or persistent directory defined by DATA_PATH. Stores nodes, edges, schema definitions, and vector embeddings backed by Delta Lake transaction logs.

DATA_PATH=/data graph/nodes/ graph/edges/ vectors/ _delta_log/ Snappy Parquet
Tier 03 // Query & GraphRAG

Serverless Control Plane

FastAPI query runtime running in-process DuckDB to execute recursive SQL CTE traversals directly against Delta tables. Exposes sub-second REST endpoints, interactive graph visualizations, and GraphRAG tools.

DuckDB C++ Engine FastAPI REST Vector Similarity Timeline Scrubbing GraphRAG Agent Tools

Ultra-Fast Recursive Graph Traversal in SQL

Rather than mandating niche graph query dialects, LineagIQ executes multi-hop graph traversals using standard recursive SQL Common Table Expressions (CTEs) evaluated directly against Delta Lake snapshots via DuckDB:

duckdb_store.py — Recursive Downstream Blast Radius Query
-- LineagIQ Recursive Downstream Lineage Traversal
WITH RECURSIVE downstream_traverse AS (
    -- Anchor: Immediate downstream dependencies of target node
    SELECT source_id, target_id, type, 1 as depth
    FROM delta_edges_view
    WHERE source_id = 'postgres.raw.raw_customers'
      AND type != 'BELONGS_TO'

    UNION ALL

    -- Recursive Step: Multi-hop graph expansion up to max_depth
    SELECT e.source_id, e.target_id, e.type, d.depth + 1
    FROM delta_edges_view e
    JOIN downstream_traverse d ON e.source_id = d.target_id
    WHERE d.depth < 5 AND e.type != 'BELONGS_TO'
)
SELECT DISTINCT source_id, target_id, type, depth
FROM downstream_traverse;

Programmatic GraphRAG Integration for AI Agents

Embed LineagIQ's lineage intelligence and blast-radius tools into agentic Python workflows (LangChain, AutoGen, LlamaIndex, or CrewAI) with minimal boilerplate:

agent_example.py — Autonomous Impact Assessment
from openai import OpenAI
from control_plane.src.agent_tools import (
    get_dataset_blast_radius,
    get_lineage_time_travel_diff
)

# 1. Retrieve Graph Context & Synthesized GraphRAG Prompt (configured via DATA_PATH)
blast_radius_prompt = get_dataset_blast_radius(
    node_id="postgres.raw.raw_customers",
    max_depth=5
)

# 2. Dispatch Context to LLM of your choice (OpenAI, Gemini, or local Ollama)
client = OpenAI()
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[
        {"role": "system", "content": "You are LineagIQ AI Agent, an expert enterprise data architect."},
        {"role": "user", "content": blast_radius_prompt}
    ]
)

print(response.choices[0].message.content)

Containerized Deployment with Unified DATA_PATH

Deploy LineagIQ effortlessly with Docker and Docker Compose. All components share a mounted directory or cloud bucket prefix orchestrated through the unified DATA_PATH environment variable:

docker-commands.sh — Production Docker Quickstart
# 1. Generate 3-Version Time-Travel Demo Dataset with Public Image
docker run --rm \
  -v $(pwd)/data:/data \
  ghcr.io/timor-dataworks/lineagiq-collection-agent:latest \
  --multiversion-demo

# 2. Launch Serverless Control Plane & Web Visualizer (Port 8000)
docker run -d --name lineagiq-control-plane \
  -p 8000:8000 \
  -e DATA_PATH=/data \
  -v $(pwd)/data:/data \
  ghcr.io/timor-dataworks/lineagiq-control-plane:latest

# Or launch entire stack via Docker Compose:
docker compose up -d control_plane

Quick Start: Run LineagIQ in 60 Seconds

Pull pre-built, production-ready multi-architecture containers directly from GitHub Container Registry. No local Python build or cloud setup required.

Control Plane & Visualizer
ghcr.io/timor-dataworks/lineagiq-control-plane:latest
linux/amd64 FastAPI REST DuckDB 1.1+ Port 8000
Stateless Collection Agent
ghcr.io/timor-dataworks/lineagiq-collection-agent:latest
linux/amd64 Delta Lake Writer PyArrow Ephemeral CLI
Step 01

Generate or Ingest Metadata

Run the ephemeral agent to generate a 3-version historical lineage evolution dataset (or point to your own dbt manifests / SQL catalogs).

Terminal
docker run --rm \
  -v $(pwd)/data:/data \
  ghcr.io/timor-dataworks/lineagiq-collection-agent:latest \
  --multiversion-demo
✓ Output written to ./data (Delta Lake tables: nodes, edges, vectors, commit logs).
Step 02

Launch Serverless Control Plane

Start the query runtime. DuckDB mounts your Delta Lake directory and immediately serves the interactive visualizer and REST APIs.

Terminal
docker run -d --name lineagiq-control-plane \
  -p 8000:8000 \
  -e DATA_PATH=/data \
  -v $(pwd)/data:/data \
  ghcr.io/timor-dataworks/lineagiq-control-plane:latest
✓ Service runs on http://localhost:8000 with sub-second response times.
⚡ All-in-One Alternative

Launch with Docker Compose

Spin up the complete pipeline (demo dataset generation + serverless control plane) with a single command:

One-Liner Quickstart
# Download docker-compose.yml and start stack
curl -sSL https://raw.githubusercontent.com/timor-dataworks/LineagIQ/main/docker-compose.yml -o docker-compose.yml
docker compose run --rm demo_generator
docker compose up -d control_plane