SoftwareLore

Category 15 software profiles

Data & AI Platforms

Cloud data warehouses, lakehouses, analytics engines and enterprise AI platforms.

Data and AI platforms help organizations turn large, messy collections of information into analysis, predictions and automated decisions. The category includes cloud data warehouses, lakehouse platforms, distributed processing engines, machine learning tooling and ontology-driven analytics software used by governments and large enterprises.

Since the arrival of generative AI, these platforms have become the foundation for building AI applications on private company data, so vendors now compete on governance, security and how easily models can be connected to trusted data. Open-source projects such as Apache Spark and MLflow sit alongside commercial platforms that package them as managed services.

Profiles

Data & AI Platforms: software on SoftwareLore

Makers

Who makes Data & AI Platforms?

Buying guide

What to look for

Questions worth answering before choosing a tool in this category.

  1. 01

    Warehouse, lakehouse or both

    Data warehouses excel at structured SQL analytics, while lakehouses also handle raw files, streaming and machine learning on open formats. Many platforms now support both patterns.

  2. 02

    Governance

    Fine-grained access control, lineage tracking and auditing are essential when sensitive data feeds dashboards, reports and AI applications.

  3. 03

    Cost model

    Most cloud data platforms separate storage from compute and bill for usage. Monitor query and cluster costs closely, because they grow quickly with adoption.

  4. 04

    AI readiness

    Evaluate how easily models, vector search and AI agents can be built on governed company data without copying that data into separate systems.

FAQ

Data & AI Platforms: frequently asked questions

What is the difference between a data warehouse and a data lakehouse?

A data warehouse stores structured, cleaned data optimized for SQL analytics and reporting. A data lakehouse combines the low-cost, flexible storage of a data lake, which holds raw files in open formats, with warehouse-style management and performance, so the same data can serve analytics and machine learning.

What is Apache Spark used for?

Apache Spark is an open-source engine for processing large datasets in parallel across clusters of computers. It is used for data engineering, SQL analytics, streaming and machine learning, and it underpins commercial platforms such as Databricks.

What is an ontology in a data platform?

In platforms such as Palantir Foundry, an ontology maps raw data to real-world objects, like customers, aircraft or shipments, along with their properties, relationships and the actions people can take on them. It lets analysts and applications work with business concepts instead of database tables.

Other categories

to move · Enter to open · Esc to close