LineagIQ unifies fragmented, multi-cloud data pipelines into an active, queryable semantic knowledge graph. Experience point-in-time schema drift analysis, instant recursive blast-radius attribution, and an ultra-lean serverless architecture that completely eliminates expensive, dedicated graph database clusters.
Data pipelines fail silently when upstream tables alter column types, rename keys, or drop fields unannounced. LineagIQ captures every structural revision and dependency shift across time through Delta Lake's ACID transactional commit logs.
Before deploying schema migrations or refactoring dbt models, LineagIQ computes downstream blast radius in milliseconds—preventing downstream data incidents and broken executive dashboards.
LineagIQ bridges graph intelligence with generative AI through Graph-Augmented Generation (GraphRAG). Ground your LLMs with deep structural lineage context to eliminate hallucinations during architectural audits and incident investigations.
The LineagIQ Collection Agent operates as a lightweight, stateless utility inside your secure VPC, CI/CD pipeline, or Kubernetes cluster. It ingests schema definitions, warehouse catalogs, audit logs, and runtime orchestrator events—transforming raw operational signals into unified graph models and semantic vector embeddings without ever touching raw customer record data.
Directly ingests compilation artifacts (manifest.json and catalog.json). Extracts data models, seeds, snapshots, source declarations, column-level documentation, tags, and tests, compiling the complete transformation DAG into unified DERIVED_FROM dependency edges.
python -m collection_agent.src.cli --dbt-manifest <path> --dbt-catalog <path>
Parses standard INFORMATION_SCHEMA.TABLES, COLUMNS, and foreign key constraint definitions. Extracts relational schemas, column data types, ordinal positions, and physical JOINS_WITH relationships across relational databases and cloud warehouses.
python -m collection_agent.src.cli --sql-schema <path_to_schema.json>
Analyzes warehouse audit histories (Snowflake QUERY_HISTORY, BigQuery INFORMATION_SCHEMA.JOBS, Databricks Query History). Mines actual user and service dataset consumption patterns (CONSUMED_BY), query execution frequencies, and implicit multi-table join predicates.
python -m collection_agent.src.cli --query-logs <path_to_logs.json>
Ingests standardized OpenLineage JSON run events emitted by workflow orchestrators and distributed compute engines. Seamlessly maps job run IDs, execution timestamps, input dataset consumption edges (CONSUMED_BY), and output production edges (PRODUCED_BY).
python -m collection_agent.src.cli --openlineage <path_to_events.json>
Legacy enterprise data catalogs mandate dedicated Neo4j, TigerGraph, or Amazon Neptune clusters—burdening teams with persistent infrastructure bills, JVM memory management, and specialized graph query skills. LineagIQ re-architects graph intelligence from first principles.
Thanks to LineagIQ's pure NumPy CSR/CSC memory layouts and Apache Arrow zero-copy ingestion, historical graph snapshots are cached directly as contiguous integer index slices. This achieves sub-millisecond (20–40 µs) lineage reachability traversals with an 18.6x memory footprint reduction—running large-scale topologies on lightweight, low-cost compute.
| Scale Tier | Graph Topology (Nodes & Edges) | In-Memory Cache Footprint | Recommended Hardware | Recursive Traversal Latency | Legacy JVM Graph Equivalent |
|---|---|---|---|---|---|
| Mid-Market | 50,000 nodes • 150,000 edges | ~1.5 MB |
0.25 vCPU, 512 MB RAM (e.g., ~$5/mo) | < 0.05 ms (20 µs) | 16 GB RAM Neo4j (~$300/mo) |
| Enterprise | 250,000 nodes • 1,200,000 edges | ~12 MB |
0.5 vCPU, 1 GB RAM (e.g., ~$10/mo) | < 0.15 ms (150 µs) | 32 GB RAM Neptune (~$850/mo) |
| Large Enterprise | 1,000,000 nodes • 5,000,000 edges | ~50 MB |
1 vCPU, 2 GB RAM (e.g., ~$20/mo) | < 0.50 ms | 64 GB RAM HA Cluster (~$2,200/mo) |
| Hyperscale Global | 5,000,000+ nodes • 25,000,000+ edges | ~250 MB |
2 vCPU, 4 GB RAM (e.g., ~$45/mo) | < 2.50 ms | 128 GB+ Distributed Cluster ($5,000+/mo) |
Designed for platform architects, data engineers, and AI developers. LineagIQ cleanly decouples edge metadata extraction, immutable table storage, and vectorized query execution.
Executes as an ephemeral batch task (Docker, Kubernetes CronJob, or GitHub Action) inside customer boundaries. Ingests metadata from files, schemas, query logs, and orchestrators.
Standard cloud object storage (S3/GCS) or enterprise storage layer. Stores lineage nodes, dependency edges, schema definitions, and vector embeddings backed by Delta Lake transaction logs.
FastAPI query runtime running a high-performance pure NumPy CSR/CSC engine to execute multi-hop lineage traversals in microseconds (< 0.05 ms) directly against Delta Lake snapshots. Exposes sub-millisecond REST endpoints, interactive graph visualizations, and GraphRAG tools.
LineagIQ transforms historical Delta Lake graph snapshots directly into in-memory Compressed Sparse Row (CSR) and Compressed Sparse Column (CSC) contiguous integer index arrays. This eliminates database query planners, virtual machines, and SQL parsing overhead to achieve sub-millisecond (< 25 µs) lineage reachability.
LineagIQ separates 1-hop structural relationships from multi-hop dataset lineage to maintain graph integrity:
fwd_all Index (Depth = 1): Indexes all edge types including BELONGS_TO and JOINS_WITH, allowing column-level nodes to resolve parent datasets.
fwd_lineage Index (Depth ≥ 2): Prunes containment edges to prevent transitive cross-dataset column leakage during deep blast-radius reachability.
targets[indptr[u]:indptr[u+1]] in continuous C-memory buffers.
Upstream root-cause investigations trace dependencies backward without reverse table scan bottlenecks:
bwd_indptr & bwd_targets) for instant root-cause tracing.
# Pure NumPy Contiguous CSR Slicing Algorithm
import numpy as np
def traverse_blast_radius(start_idx: int, max_depth: int, fwd_all, fwd_lineage):
visited = {start_idx}
queue = [(start_idx, 0)]
impacted_nodes = []
while queue:
curr_idx, depth = queue.pop(0)
if depth >= max_depth:
continue
# Select dual CSR index: fwd_all for depth 1, fwd_lineage for depth >= 2
indptr, targets = (fwd_all if depth == 0 else fwd_lineage)
# Zero-overhead contiguous slice: targets[indptr[curr_idx]:indptr[curr_idx + 1]]
start_pos = indptr[curr_idx]
end_pos = indptr[curr_idx + 1]
neighbors = targets[start_pos:end_pos]
for neighbor in neighbors:
if neighbor not in visited:
visited.add(neighbor)
queue.append((neighbor, depth + 1))
impacted_nodes.append(neighbor)
return impacted_nodes # Executed in ~0.02 ms (20 microseconds)
LineagIQ generates high-fidelity semantic representations in-process using an INT8-quantized ONNX embedder (bge-small-en-v1.5). Raw customer data never leaves private network perimeters, and vector artifacts are persisted with full ACID versioning in Delta Lake.
Standard text embeddings fail on relational schema tokens. LineagIQ formats specialized prompt representations tailored to each node ontology:
dataset table {name}: {description} with columns {cols}
column {name} of type {type} in dataset {dataset}: {desc}
pipeline ETL transformation {name}: {desc}
business term governance {name}: {definition}
Dense embeddings are stored alongside graph metadata to support point-in-time semantic search:
||v||₂ = 1.0), simplifying cosine similarity to a BLAS dot product.
vectors/ Delta tables, maintaining atomic parity with graph node commits.
as_of) to discover deleted or renamed historical assets.
# In-process quantized embedding generation & L2 unit normalization
from core.embedder import LocalEmbedder
from core.models import Node, NodeType
# 1. Initialize quantized INT8 ONNX embedder (zero external API latency)
embedder = LocalEmbedder()
# 2. Node with domain properties
node = Node(
id="analytics.mart_orders",
type=NodeType.DATASET,
name="mart_orders",
description="Consolidated customer orders mart for revenue analysis"
)
# 3. Generate domain-specific representation and normalized 384-dim vector
semantic_text = embedder.format_node_text(node)
embedding_vector = embedder.embed_text(semantic_text)
# Norm verification: np.linalg.norm(embedding_vector) == 1.0
print(f"Generated {len(embedding_vector)}-dim normalized vector: {embedding_vector[:4]}...")
LineagIQ bridges the gap between semantic catalog discovery and structural knowledge graphs. By evaluating dense vector similarities via hardware-accelerated BLAS matrix operations and overlaying multi-hop topological reachability, LineagIQ synthesizes rich GraphRAG prompts for autonomous LLM agents.
Bypasses bloated vector database daemon architectures through contiguous in-memory linear algebra:
(N, 384) matrix in RAM via Apache Arrow zero-copy buffers.
scores = matrix @ query_vector.
np.argpartition(-scores, top_k) extracts highest scoring matches in linear time, avoiding O(N log N) sorting bottlenecks.
Vector search finds the root; graph traversal discovers the operational blast radius:
→), fully compatible with OpenAI, Gemini, and Ollama.
# GraphRAG Context Synthesis: Combining Vector Search + CSR Blast Radius
from control_plane.src.query_engine import QueryEngine
from control_plane.src.prompt_synthesizer import PromptSynthesizer
# 1. Initialize Unified Query Engine over Delta Lake
engine = QueryEngine(data_path="/data/tenants/default/")
# 2. Hybrid Step A: Semantic Vector Discovery via BLAS Dot-Product
candidates = engine.vector_store.search_similar(query="customer churn revenue", top_k=1)
target_node_id = candidates[0]["id"] # "analytics.marts.fct_churn"
# 3. Hybrid Step B: Microsecond CSR Multi-Hop Traversal (5 Hops)
traversal_result = engine.graph_store.traverse_blast_radius(
node_id=target_node_id,
max_depth=5
)
# 4. GraphRAG Context Synthesis: Format structured prompt for AI Agent
synthesizer = PromptSynthesizer()
llm_prompt = synthesizer.build_blast_radius_prompt(traversal_result)
# Send to OpenAI, Gemini, or local Ollama instance (http://localhost:11434)
Embed LineagIQ's lineage intelligence and blast-radius tools into agentic Python workflows (LangChain, AutoGen, LlamaIndex, or CrewAI) with minimal boilerplate:
from openai import OpenAI
from control_plane.src.agent_tools import (
get_dataset_blast_radius,
get_lineage_time_travel_diff
)
# 1. Retrieve Graph Context & Synthesized GraphRAG Prompt
blast_radius_prompt = get_dataset_blast_radius(
node_id="postgres.raw.raw_customers",
max_depth=5
)
# 2. Dispatch Context to LLM of your choice (OpenAI, Gemini, or local Ollama)
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{"role": "system", "content": "You are LineagIQ AI Agent, an expert enterprise data architect."},
{"role": "user", "content": blast_radius_prompt}
]
)
print(response.choices[0].message.content)
Pull pre-built, production-ready multi-architecture containers directly from GitHub Container Registry. No local Python build or cloud setup required.
ghcr.io/timor-dataworks/lineagiq-control-plane:latest
ghcr.io/timor-dataworks/lineagiq-collection-agent:latest
Run the ephemeral agent to generate a 3-version historical lineage evolution dataset (or point to your own dbt manifests / SQL catalogs).
docker run --rm \
-v $(pwd)/data:/data \
ghcr.io/timor-dataworks/lineagiq-collection-agent:latest \
--multiversion-demo
./data (Delta Lake tables: nodes, edges, vectors, commit logs).
Start the query runtime. LineagIQ mounts your Delta Lake directory and immediately serves the interactive visualizer and REST APIs.
docker run -d --name lineagiq-control-plane \
-p 8000:8000 \
-e DATA_PATH=/data \
-v $(pwd)/data:/data \
ghcr.io/timor-dataworks/lineagiq-control-plane:latest
http://localhost:8000 with sub-second response times.
Spin up the complete pipeline (demo dataset generation + serverless control plane) with a single command:
# Download docker-compose.yml and start stack
curl -sSL https://raw.githubusercontent.com/timor-dataworks/LineagIQ/main/docker-compose.yml \
-o docker-compose.yml
docker compose run --rm demo_generator
docker compose up -d control_plane