Speakers

Speaker "Maximo Gurmendez" Details

Name :
maximo gurmendez
Company :
Title :
Lead Engineer
Topic :

Building a petabyte scale machine learning engine using Apache Spark

Abstract :

The central premise of DataXu is to apply data science to better marketing. At its core, is the Real Time Bidding Platform that processes 2 Petabytes of data per day and responds to ad auctions at a rate of 2.1 million requests per second across 5 different continents. Serving on top of this platform is Dataxu’s analytics engine that gives their clients insightful analytics reports addressed towards client marketing business questions. Some common requirements for both these platforms are the ability to do real-time processing, scalable machine learning, and ad-hoc analytics. This talk will showcase DataXu’s successful use-cases of using the Apache Spark framework to address all of the above challenges while maintaining its agility and rapid prototyping strengths to take a product from initial R&D phase to full production. The team will share their best practices and highlight the steps of large scale Spark ETL processing, model testing, all the way through to interactive analytics.

Profile :
Maximo holds a Masters degree in Computer Science / Artificial Intelligence from Northeastern University where he attended as a Fulbright Scholar. Since 2009 he has been working with DataXu as a lead engineer, tackling the challenge of machine learning over large large data sets. He’s also the founder of MDATALABS (data science & engineering consultancy) and a Big Data Science professor at the School of Engineering, University of Montevideo.
x

Get latest updates of Global Data Science Conference
sent to your inbox.

Weekly insight from industry insiders.
Plus exclusive content and offers.