One platform for the whole data lifecycle
Most teams stitch together an ingestion tool, a transformation tool, a modelling layer, an orchestrator, a BI tool and a catalogue, then spend their time keeping the seams from splitting. DataLens is those seven stages in one place, with an AI assistant that has read your data and can act on it.
The seven stages
Each one is a working surface in the product, not a roadmap item.
📥 Data
Land rows from files, databases, APIs, CRMs, lakes and document stores.
🔄 Prepare
Fix nulls, duplicates, types and names - with the AI proposing the fixes.
🏗 Model
Reverse-engineer a source model, design a target, and map between them.
⚡ Orchestrate
Build the flow on a canvas, run it, version it, promote it.
📊 Analyse
Charts, pivots and scenario simulation over the data you just prepared.
🚚 Deliver
Publish governed data products and access-controlled shares.
🛡 Govern
Profiling, quality rules, lineage, PII detection and AI cost observability.
What makes it different
The AI has read your data
Ari is not a chat box bolted to the side. It profiles what you loaded, proposes the transforms, explains what it found, and applies the change when you accept it.
No pipeline code required
Flows are built on a canvas - sources, transforms, joins, gates, AI steps - and run from the same screen you designed them on.
Governance is not a separate product
Profiling, quality rules, PII detection and lineage sit on the same datasets you are transforming, so the catalogue cannot drift from the pipeline.
Degrades instead of failing
Every optional capability probes for what it needs and switches itself off with a message when it is absent, rather than taking the session down.
Your models, your keys
Bring your own AI model credentials. They are held as a connection like any other source, and the platform reports what every AI step cost.
Reversible by design
Transforms are recorded per dataset, working copies leave the original intact, and a flow can be reset to its base at any point.
A first session, end to end
Land the data
Drop in a CSV or Excel file, or connect a database, REST API, SFTP drop, CRM or data lake. It is profiled the moment it arrives - types, distributions, null rates, candidate keys.
Ask what is wrong with it
Ari reports the quality problems it found and proposes fixes. You accept the ones you want; each becomes a recorded step against that dataset.
Give it a shape
Reverse-engineer a source model from what arrived, design the target you actually want, and map between them field by field.
Make it repeat
Turn the work into a flow on the canvas, run it, and promote it through environments as a versioned release.
Answer the question
Chart it, pivot it, or run a scenario against it - then publish the result as a governed data product or an access-controlled share.
Start from where your data already lives
Nine ways in, all of which land in the same catalogue.
CSV & Excel upload
Drag a file in and start working. Profiled on arrival, no schema declared up front.
Database
Read from your existing relational database and pull tables in as datasets.
REST API
Point at an endpoint, map the response shape, and land it as rows.
FTP & SFTP
Collect the drops that still arrive as files on a server.
CDC sync
Change data capture, so the dataset follows the source instead of ageing.
CRM
Bring customer records across without exporting them to a spreadsheet first.
Data lake
Read from - and write back to - object storage you already own.
Documents
PDFs and documents stored whole, for retrieval and generation rather than rows.
AI model providers
Your own model credentials, held as a connection like any other source.
And when the next step is a model
The machine learning layer is in build. What it will enforce is written down now rather than afterwards.
Machine learning
A model registry, per-environment approval, features, splits and recorded runs - on the datasets you already govern.
Model trust and approval
A model version runs in an environment because somebody approved it there, and the approval expires.
Feature store
Versioned features that carry the sensitivity of the columns they were built from.
Common questions
Do I need to write code to use DataLens?
No. Transforms, models, mappings and flows are all built through the interface, and the AI assistant can propose and apply most of the routine work. Nothing stops you writing SQL or a script node where that is the clearer answer - the flow canvas has script nodes for exactly that.
Where does my data go?
Data lands in the platform for the session you are working in, and can be routed to object storage you own rather than ours. AI model calls go to the provider whose credentials you supplied.
Does DataLens replace my data warehouse?
No, and it is not trying to. It sits in front of one - landing, cleaning, modelling and governing data on the way in, and publishing governed products on the way out.
Can I run machine learning models on this data?
DataLens does not train or serve models - it governs them. The model registry, per-environment approval, the feature store, reproducible splits, run history and egress control are implemented, and the Data Science tab reads all of it. The model itself runs where it already runs: a REST endpoint, Azure ML, SageMaker, Vertex AI, Databricks, Fabric, Snowflake or MLflow.
Is DataLens available now?
It is in private beta. You can request access, and beta users work with the full lifecycle rather than a cut-down trial.
See it on your own data
DataLens is in private beta. Bring a file, a database or an API and work through the whole lifecycle in one sitting.
Request beta access