Data Scientist
Required skills for this role
OracleSQLPythonAWSKafkaDevOpsCI/CDMongoDBPostgreSQLMySQLHadoopSparkSnowflake
About this role
Job Description
Data Architect / Data Scientist
Location :- Chennai
Experience :- 6–9 Years
Choosing Capgemini means choosing a place where you'll be empowered to shape your career, supported by a collaborative global community, and inspired to reimagine what's possible. Join us in designing next-generation data platforms that enable advanced analytics, machine learning, and real-time decision-making for global enterprises.
Your Role
As a Data Architect / Data Scientist, you will be responsible for designing scalable, secure, and high-performance data platforms. You will work closely with engineering teams and business stakeholders to build modern data ecosystems that support analytics, AI/ML initiatives, and real-time data processing.
In this role, you will:
• Design end-to-end data architectures across cloud and hybrid environments.
• Build scalable data pipelines, data lakes, and modern data platforms.
• Architect ingestion frameworks for IoT and sensor data, particularly for Digital Twin platforms.
• Develop and optimize data models for structured, semi-structured, and unstructured data.
• Implement real-time data ingestion using streaming and messaging frameworks.
• Collaborate with cross-functional teams to align data strategies with business goals.
• Define and enforce data governance, security, quality, and compliance standards.Your Profile
• 6–9 years of progressive experience in data engineering and data architecture.
• Strong experience designing scalable and secure data platforms.
• Hands-on expertise with AWS data services and modern data platforms.Core Skills & Expertise
• Cloud Expertise: AWS (Kinesis, EMR, Glue, RDS, Athena, Redshift, Lambda, S3)
• Data Platform Design: Snowflake, scalable data pipelines, data lake architectures
• Digital Twin Integration: IoT data ingestion, sensor data modeling, real-time systems
• Data Modeling: Conceptual, logical, and physical modeling across data types
• Programming & Querying: Python, PySpark, Advanced SQL
• Big Data Ecosystem: Hadoop, Spark, Kafka
• Database Management: SQL Server, MySQL, PostgreSQL, Oracle, MongoDB, Cassandra
• Data Governance & Quality: Lineage, metadata management, quality frameworks
• DevOps & CI/CD: Deployment automation and pipeline management for data platforms
• Streaming & Queuing: Real-time ingestion architectures
• Time Series Data: Design for ingestion, storage, and analysis of time-based dataWhat You'll Love About Working Here
• Opportunity to work on cutting-edge data platforms and AI-driven solutions.
• Collaborative and innovative work culture.
• Continuous learning opportunities in cloud and data technologies.About Us
Capgemini is a global business and technology transformation partner, helping organizations accelerate their digital and data transformation journeys. With strong expertise in cloud, data, AI, and engineering, Capgemini delivers end-to-end services that empower enterprises to innovate and scale efficiently.
At Capgemini Engineering, the world leader in engineering services, we bring together a global team of engineers, scientists, and architects to help the world's most innovative companies unleash their potential. From autonomous cars to life-saving robots, our digital and software technology experts think outside the box as they provide unique R&D and engineering services across all industries. Join us for a career full of opportunities. Where you can make a difference. Where no two days are the same.