[ Tech stack ]

Apache Spark

The distributed analytics engine for processing terabytes.

Spark processes massive data volumes in memory across a distributed cluster. Its APIs (DataFrame, SQL, Streaming, MLlib) cover batch, stream and machine learning. Paired with Delta Lake, it's the foundation of modern lakehouses.

[ Why Apache Spark at Dexon ]

What this technology does well, and why we use it.

Typical usage: Large-volume ETL, lakehouse, distributed ML training.

  • 01

    Distributed in-memory processing: orders of magnitude vs Hadoop.

  • 02

    Python, Scala, SQL, R APIs: mixed teams accommodated.

  • 03

    Structured Streaming for near-real-time.

  • 04

    Databricks, EMR, Dataproc: managed on all three clouds.

[ Where we use Apache Spark ]

One area of expertise involved.

Apache Spark is part of the technical composition of our engagements on the following areas of expertise. Click to discover the scope.

[ Complementary technologies ]

The building blocks we often mobilise alongside.

A stack rarely exists alone. Here are the technologies Dexon most often pairs with this one, through pipeline habits, usage similarity or internal mastery. Click on a building block to see its scope.

[ Reassurance ]

0+
custom projects delivered
30
engineers, designers, project managers
80 %
from top French schools
24 h
average reply time

[ They trust us ]

More than 100 French and European companies trust us

[ Press ]

They talk about us.

Application and data division.

Nationwide coverage by BFM Business, Le Figaro, Challenges, La Tribune and CNews. An outside reading of our work and our innovations.

Start a project

A project with Apache Spark?

Describe your Apache Spark project. Our team comes back within 24 h with free technical scoping. No commitment.

  • Reply within 24 h by a specialized consultant.
  • Free technical scoping, with no fees.
  • No commitment, your data stays confidential.
01 73 22 43 74contact@dexon.fr

Present inParisLyonMarseilleNiceGenève