DPF (Data Prompt Forge) simplifies data ingestion and transformation through plain natural language prompts. Describe the data you have and the shape you want â the platform's AI infers the schema and generates the parsing and transformation logic for you in seconds, not weeks. No expensive ETL tools, no hand-coded data pipelines.
It's an open data management platform built on Apache Iceberg. Your data lands in open, vendor-neutral tables with schema evolution, time-travel, and rollback â never locked into a proprietary warehouse. Query it instantly with the built-in query engine, or follow our integration guides to bring your own compute: Spark, Trino, Databricks, Redshift, and more.
Everything is API-first and built to integrate with AI: every action in the web portal is also a REST API call or available through our MCP server, so AI agents can onboard data sources, run transformations, and query results without a human in the loop.
Describe your data and the result you want in plain language. Our AI â augmented with proprietary integration templates and industry best practices â infers the schema and generates production-ready parsing and transformation code automatically. No hand-coding, no separate ETL tool.
Traditional ETL projects take weeks of development and testing. DPF delivers working data pipelines in minutes â from upload to queryable data in open table format.
Traditional stacks require an ETL tool (Fivetran, Informatica) plus a warehouse (Snowflake, BigQuery). DPF bundles transformation and managed Iceberg storage into a single usage-based subscription. Connect your own query engine when you need heavy analytics.
CSV, fixed-width, JSON, XML, Parquet, Excel, and proprietary vendor formats. AI detects the structure â multi-record types, nested data, custom delimiters â and handles it all.
An open data management platform built on Apache Iceberg â open tables with schema evolution, time-travel, and rollback, with no proprietary lock-in. Query instantly with the built-in engine, or follow our integration guides to bring your own compute: Spark, Trino, Databricks, or Redshift.
Every data load captures Iceberg snapshots. Made a mistake? Roll back to the exact pre-load state with a single API call. Full audit trail included.
JWT authentication, workspace-level access control, encrypted credentials, and complete audit logging. Your data stays in your AWS account.
Every action in DPF is API-driven â the web portal itself is just one client of the same REST API. Automate your entire data integration pipeline programmatically, drive it from our MCP server, or use the web portal.
Shared workspaces with granular permissions. Invite team members, share data specs, and collaborate on transformations without stepping on each other.
One-time setup process to define parsing rules, transformation logic, and mapping specifications that are reused for every execution.
Automated execution pipeline that parses input files, validates data, applies transformations, and loads results to open table formats.
Comprehensive, API-driven platform â every capability (integration management, job execution, credential handling, and data consumption) is exposed as a REST API call, not just bolted on after the fact.
DPF is built to integrate with AI from the ground up. Our Model Context Protocol (MCP) server lets AI agents autonomously set up data integrations, manage transformations, and make data available â using the same API-driven actions a human would take in the portal.
User-friendly interface for integration setup, file uploads, job monitoring, and results visualization with team collaboration features.
Share integrations, collaborate on transformations, manage permissions, and work together seamlessly across teams.
Explore our API documentation to learn how to integrate the Data Integration Service into your application, or log in to access the web portal.