Millions of rows
Datasets and transformations handle millions of rows of data and calculations, lookups, calculated columns, filters, and aggregates all run at production scale, on schedule, every time.
From the first data source to a chained, scheduled, lineage-traced downstream dataset, each step is auditable and isolated.
01 · Bring in data
Mud Cake accepts data from multiple sources. Upload spreadsheet files (.xlsx), extract structured information from business documents (.pdf, .docx), or connect to your existing SQL data tables. Each source lands as a governed, inspectable dataset ready for transformation.
02 · Extract & validate
For document sources, the worker parses each file and runs extraction and validation through AI-backed services using your configured prompts and thresholds. Extracted data objects carry confidence, validation scores, and review status, so nothing becomes a dataset until it has been checked.
03 · Catalog & organize
Work centers on datasets. Browse and filter the catalog, open a dataset for schema, preview, activity, and lineage, and organize datasets into nested folders. Folders are logical organization only, they never change physical storage or tenant scope.
Invoices · 2026
spreadsheet-backed
Revenue Summary
transformation output
Vendor Lookup
transformation output
Add · Duplicate · Delete · Reorder, every repeatable type, in explicit order.
04 · Transform, no SQL
Build transformations from six structured operation types. Lookups can join your data against other Mud Cake datasets or your existing SQL tables, combining spreadsheet data with large-scale enterprise datasets. The backend generates trusted SQL from validated identifiers and injects tenant and customer predicates. Users never author SQL, and cross-tenant selection is prohibited.
05 · Automate & chain
Schedule transformations hourly, daily, weekly, or monthly. When a run succeeds, Mud Cake durably queues downstream transforms automatically. Run history records Manual, Scheduled, and Dependency origin with bounded direct-upstream causality.
Invoices · 2026
Manual · succeeded
Revenue Summary
Dependency · succeeded
Vendor Lookup
Dependency · succeeded
06 · Govern & isolate
Customer is the SaaS ownership boundary; tenant is the workspace and data-isolation boundary. Every dataset, transformation, schedule, folder, and execution stays tenant-scoped, and backend authorization, not UI filtering or a parsed physical name, is the authoritative boundary.
Read the full governance modelScale & configuration
Mud Cake handles millions of rows of data and calculations without breaking a sweat, and getting there is easy, not engineering.
Datasets and transformations handle millions of rows of data and calculations, lookups, calculated columns, filters, and aggregates all run at production scale, on schedule, every time.
Configure processes, extraction prompts, validation thresholds, and transformations through a guided UI, no code, no SQL, no scripts. If your team can use a spreadsheet, they can run Mud Cake.
A customer can own multiple tenants, map them to divisions, regions, or business units. Each tenant is independently governed and fully isolated, with its own users, datasets, and audit trail.
Marry data from your spreadsheets and documents with your existing SQL data tables. Look up and join against large-scale enterprise datasets, and run calculations across all of them at scale.
Request a walkthrough tailored to your data and your tenants.