Megaphone

Get a Free 30-Day Proof of Concept — we cover the cost.

Home / Solutions / Data Architects
Keep the engine and deployment choices reversible

DataFuseAI for Data Architects

The engine you standardised on two years ago, now quietly deciding what every new workload is allowed to be — and the third team asking for its own workspace, which becomes a governance exception nobody is tracking.

DataFuseAI keeps the engine decision reversible. Point a workload at the native engine, a Databricks workspace, or your own Spark cluster through Apache Livy; give each team its own tenant workspace; and run the platform managed, private-hosted, or on-premise. One RBAC and audit-log model covers all of it.

Day one: create a tenant workspace, register an engine profile, and run one pipeline against it.

14-day free trial · No credit card required · Managed, private-hosted, or on-premise
SOUND FAMILIAR?

Where the Architecture Stops Being Yours

The architecture rarely fails at the design. It erodes at the edges — an engine choice that hardens into a constraint, a standard that survives three teams and not the fourth, a regulated workload that quietly leaves the platform, and isolation held together by an agreement rather than by a permission. Those four are where the reference architecture and the running one part ways.

The engine you standardised on before the workloads existed

The decision was right for what you had. Two years on, a workload wants a Databricks cluster the platform was never wired for, and the choice is to bend the workload or stand up a second platform beside the first. An engine chosen once quietly becomes a constraint on everything built after it, and the cost lands on whoever proposes the next workload. Lock-in is easier to see when you ask where the data and the compute actually sit.

The standard that holds until the fourth team

A naming convention, a currency rule, a definition of "active customer" — agreed in a document, implemented properly in the first pipeline, approximated in the second, and re-derived from scratch in somebody's workbook by the fourth. A definition enforced by convention lasts exactly as long as everyone has time to honour it. By the time two reports disagree, tracing which one is wrong costs more than writing the rule did.

The workload that is not allowed to leave the building

Residency rules, an on-premise requirement, an environment with no outbound connectivity — and a cloud-only platform that cannot go there. So the regulated workload gets its own quiet toolchain: scripts on a box inside the perimeter, outside the platform everyone else is governed by. The workload with the strictest requirements ends up on the tooling with the least oversight. Worth settling early is what an on-premise deployment actually has to satisfy.

Isolation that exists by agreement rather than by permission

Two teams share a workspace because splitting it was never anyone's sprint. They share a service account, a set of credentials, and an understanding that nobody touches the other team's tables. "Nobody else uses that account" is not an access control — and the gap only becomes a finding when a review asks who could have reached what, and the honest answer is anyone holding the password.

None of that was in the architecture you designed.

HOW DATAFUSEAI HELPS

What the Platform Decides, and What Stays Reversible

Three decisions stay reversible: which engine runs a workload, which tenant a team works in, and where the platform itself is deployed. Each of the three is configuration rather than a rebuild, and changing one does not change the other two. What does not move is the layer underneath — one RBAC and audit-log model applies across every tenant, every engine, and every deployment option.

Choose the engine per workload, not per platform

Instead of one engine locking the architecture in, choose per workload and keep one governance model. An Engine Profile points execution at the DataFuseAI native engine, a Databricks workspace you already operate, or a Spark cluster reached through Apache Livy. The pipeline definition is held separately from the runtime, so moving a workload onto different compute is a change of profile rather than a rebuild.

A tenant boundary the permission model enforces

Each tenant operates as an isolated workspace with its own users, groups, and configuration, and a SuperAdmin manages company-level settings, access, and module visibility from one place. Role-based access control and action-level permissions are part of the core architecture rather than a layer added afterwards, which is what makes the isolation dependable rather than conventional. That is how the platform is put together underneath the workspaces.

The same platform wherever the data is allowed to live

Managed, private-hosted, and on-premise, with the same platform features in each — and on-premise runs locally with no external connectivity requirement, so the environment and the data path stay inside your perimeter. The deployment decision is independent of where execution happens, which is configured separately through engine profiles. Role-based access and the audit log come from that same core architecture in every model, so choosing where to run it does not change the governance story.

One definition, reused rather than re-derived

A field definition — a unit conversion, a currency rule, a derived status — is expressed once in a pipeline and reused by every job downstream, instead of being re-derived in each team's workbook. Where teams disagree about a definition, the disagreement becomes visible in one place rather than hidden in several. That is the practical case for holding the transformation rules in one layer.

KEY BENEFITS

What the Architecture Keeps Open

Each of these traces to a mechanism you can inspect before committing a workload: an engine profile you can repoint, one permission model spanning every tenant, connection profiles that read sources where they already are, and a transformation layer that holds the definitions. None of it depends on a migration finishing first. That is the part an architect can check rather than take on trust.

Move a workload onto different compute by changing its engine profile — the pipeline definition does not have to change with it.

One RBAC and audit-log model covers every tenant workspace — there is no separate access policy per team to keep in step.

Sources are read in place through connection profiles — nothing has to migrate into the platform to be governed by it.

A field definition is expressed once and reused by every job downstream, instead of being re-derived in each team's workbook.

Where the platform runs and where the work executes are separate choices — the deployment model and the engine profile are configured independently.

CUSTOMER QUOTES

From Teams Already Running on DataFuseAI

These are DataFuseAI customers describing the platform in their own words. We have not filtered them by role or by job title.

"We used to manage dozens of scripts across servers. Any change — a driver update, credential change, or schema tweak — broke something. Moving to DataFuseAI meant pipelines, jobs, and connection profiles now live in one place. It didn't just eliminate complexity — it also made things visible and manageable, which is a huge difference."

"Scheduling used to be the weak link. A cron misfire or a missed run meant downstream reports were wrong and nobody noticed. With jobs, alerts, and the dashboard, we see failures early, and the team trusts the numbers again. There are still custom workloads we run on Databricks, but DataFuseAI coordinates the day-to-day work cleanly."

"Our analysts were constantly blocked waiting for the engineering team. With visual pipelines and saved queries, they can build and view most of what they need themselves."

"We deploy in environments with strict compliance requirements. For us, the tenant model and offline / on-premise option mattered. We can isolate data, control permissions at a fine level, and still give teams a single place to operate pipelines and queries. It's not 'magic,' but it's practical — and audit conversations are easier now."

FAQ

Frequently Asked Questions for Data Architects

Between your source systems and the destinations you already report from. It reads databases, cloud warehouses, object storage, files, and REST APIs as sources, applies transformation logic, and writes to the sink you choose. Nothing becomes the system of record that was not already — the source systems keep that role.

No. Sources are read in place through connection profiles, and outputs land in destinations you nominate. That is a deliberate architectural property: the platform is a processing and orchestration layer rather than a storage tier, so it can be introduced beside the existing estate rather than in front of it. Where the data sits matters more to vendor lock-in than which tool builds the pipeline does.

Each tenant operates as an isolated workspace with its own users, groups, and configuration, and a SuperAdmin manages company-level settings, access, and module visibility. Role-based access control and action-level permissions are part of the core architecture rather than a layer added afterwards, which is what makes the isolation dependable rather than conventional.

Give them different engine profiles. A profile points execution at the DataFuseAI native engine, a Databricks workspace you already operate, or a Spark cluster reached through Apache Livy, and each pipeline runs against the profile it is assigned. The two teams then share one platform, one permission model, and one audit trail while their work executes on separate compute, because the pipeline definition is held separately from the runtime.

Yes, and it is the main argument for centralising the work. A field definition — a unit conversion, a currency rule, a derived status — is expressed once in a pipeline and reused by every job downstream, instead of being re-derived in each team's workbook. Where teams disagree about a definition, the disagreement becomes visible in one place rather than hidden in several.

By producing the table in one place rather than policing it in several. The pipeline that writes an output defines its columns, their types, and their names, so every consumer of that table inherits the same shape instead of agreeing to it separately. A profiling step then reports the actual data type, distinct and null counts, and distribution per column on each run, so a drift away from the agreed shape shows up as a change in those numbers. Where a convention still has to be agreed rather than enforced, it at least lives in a pipeline definition that someone who did not write it can read.

Managed, private-hosted, and on-premise, with the same platform features in each. On-premise runs locally with no external connectivity requirement, so the environment and the data path stay inside your perimeter. The deployment decision is independent of where execution happens, which is configured separately through engine profiles. The trade-offs between managed, private-hosted and on-premise are worth setting out before the first environment is built.

Yes. Audit logs record changes with before-and-after detail, so a configuration or permission change can be read back rather than reconstructed from memory. Access itself is assigned to groups rather than to individuals, and action-level permissions separate running a pipeline from editing it, which is what keeps the log meaningful about who was able to do what. Job run history sits alongside it, so the record covers both what was changed and what was executed.

Batch. From connection profiles through to jobs, the architecture is built for repeatable, stable, monitored batch workflows with recorded run history. That is a design decision worth knowing early: it fits scheduled reconciliation, consolidation, and reporting workloads rather than continuous event processing. Where the line falls between batch and real-time streaming is worth deciding before the architecture assumes one.

Your next workload, on the engine you choose

Register an Engine Profile, give each team its own tenant workspace, and run the platform managed, private-hosted, or on-premise — the architecture decisions stay yours to revisit, and the RBAC and audit-log model does not change when you do.