Features

Every capability is a governed, tenant-scoped surface.

From the first data source to a chained, scheduled, lineage-traced downstream dataset, each step is auditable and isolated.

01 · Bring in data

Multiple ways to get data into Mud Cake.

Mud Cake accepts data from multiple sources. Upload spreadsheet files (.xlsx), extract structured information from business documents (.pdf, .docx), or connect to your existing SQL data tables. Each source lands as a governed, inspectable dataset ready for transformation.

  • Spreadsheet upload with durable storage and versioning
  • AI-assisted extraction from business documents with configurable prompts
  • Connect to existing SQL tables for lookups and joins
  • Queue-backed processing for durable, decoupled handling
upload to store to queuequeued
quarterly_data.xlsx2.4 MB
SQL: customers tableconnected
vendor_invoice.pdf412 KB
dataobject · validation98% confidence
invoice_noINV-8842
amount12,480.00
due_date2026-09-30
statusapproved

02 · Extract & validate

AI-assisted extraction with human review.

For document sources, the worker parses each file and runs extraction and validation through AI-backed services using your configured prompts and thresholds. Extracted data objects carry confidence, validation scores, and review status, so nothing becomes a dataset until it has been checked.

  • Configurable extraction and validation prompts per step
  • Confidence scoring and structured validation
  • Review-and-approval workflows before materialization

03 · Catalog & organize

Datasets are the primary asset, browseable, inspectable, organized.

Work centers on datasets. Browse and filter the catalog, open a dataset for schema, preview, activity, and lineage, and organize datasets into nested folders. Folders are logical organization only, they never change physical storage or tenant scope.

  • Schema contract, preview, activity, and lineage per dataset
  • Nested folders with folder-aware dataset selection
  • Separate producing-transformation and downstream-transformation surfaces
dataset catalog

Invoices · 2026

spreadsheet-backed

lineage

Revenue Summary

transformation output

lineage

Vendor Lookup

transformation output

lineage
transformation builderno user SQL
selectChoose columns
renameRename outputs
calculatedArithmetic columns
lookupLeft join another dataset
filterComparison / set filters
aggregateGroup & summarize

Add · Duplicate · Delete · Reorder, every repeatable type, in explicit order.

04 · Transform, no SQL

Repeatable, versioned recipes, without writing SQL.

Build transformations from six structured operation types. Lookups can join your data against other Mud Cake datasets or your existing SQL tables, combining spreadsheet data with large-scale enterprise datasets. The backend generates trusted SQL from validated identifiers and injects tenant and customer predicates. Users never author SQL, and cross-tenant selection is prohibited.

  • Draft, Publish, Run lifecycle with versioned definitions
  • Repeatable ordered operations: each step consumes the prior step's working schema
  • Validate and Preview against live storage before publishing
  • Clone a saved draft or published version into a new transformation
  • Look up against other datasets or your existing SQL tables

05 · Automate & chain

Schedule runs. Chain downstream. Track causality.

Schedule transformations hourly, daily, weekly, or monthly. When a run succeeds, Mud Cake durably queues downstream transforms automatically. Run history records Manual, Scheduled, and Dependency origin with bounded direct-upstream causality.

  • Successful runs durably queue downstream transforms automatically
  • Manual runs can propagate downstream or run transformation-only
  • Run history records Manual, Scheduled, and Dependency origin
run history

Invoices · 2026

Manual · succeeded

root

Revenue Summary

Dependency · succeeded

triggered

Vendor Lookup

Dependency · succeeded

triggered
authorization model
Backend authorization is authoritative
Tenant & customer predicates, server-injected
No cross-tenant selection, lookup, or dependency
Full dataset-to-dataset lineage

06 · Govern & isolate

Multi-tenant isolation, enforced server-side.

Customer is the SaaS ownership boundary; tenant is the workspace and data-isolation boundary. Every dataset, transformation, schedule, folder, and execution stays tenant-scoped, and backend authorization, not UI filtering or a parsed physical name, is the authoritative boundary.

Read the full governance model

Scale & configuration

Millions of rows. Easy configuration. Flexible tenant architecture.

Mud Cake handles millions of rows of data and calculations without breaking a sweat, and getting there is easy, not engineering.

Scale

Millions of rows

Datasets and transformations handle millions of rows of data and calculations, lookups, calculated columns, filters, and aggregates all run at production scale, on schedule, every time.

Configuration

Easy, no-code setup

Configure processes, extraction prompts, validation thresholds, and transformations through a guided UI, no code, no SQL, no scripts. If your team can use a spreadsheet, they can run Mud Cake.

Architecture

Flexible multi-tenant

A customer can own multiple tenants, map them to divisions, regions, or business units. Each tenant is independently governed and fully isolated, with its own users, datasets, and audit trail.

Integration

Connect your SQL

Marry data from your spreadsheets and documents with your existing SQL data tables. Look up and join against large-scale enterprise datasets, and run calculations across all of them at scale.

See the full platform in action.

Request a walkthrough tailored to your data and your tenants.