Overview In this on-site role in Austin, you will manage and optimize Hadoop, Big Data, and Kafka clusters in production, ensuring scalable and secure data processing. You will work with cloud and on-prem environments using AWS EMR, MSK, and GCP, collaborating with cross-functional teams to maintain high-performance data pipelines. You'll apply strong problem-solving and architectural skills to meet reliability and security needs, with a focus on observability and operational excellence. This is an hands-on role for someone who shapes data infrastructure at scale and partners across the organization.
Responsibilities- Manage and optimize Hadoop, Big Data and Kafka clusters in production
- Leverage AWS EMR, MSK, and GCP across cloud and on-prem environments
- Design and implement high-performance system architectures
- Ensure data security and privacy
- Enhance observability using Grafana, Opera, and Splunk
- Maintain Linux-based systems including networking, CPU, memory, and storage
- Collaborate with cross-functional teams to resolve issues and improve data pipelines
Key requirements- Experience managing and optimizing Hadoop, Big Data and Kafka clusters in production
- Experience with AWS EMR, MSK and GCP
- Proficiency in Python, Java or full stack development
- Familiarity with Spark, Kafka, HDFS, MapReduce
- Strong knowledge of system architecture for high-performance computing
- Understanding of data security and privacy
- Excellent problem-solving and troubleshooting skills
- Strong communication and collaboration skills
- Observability tooling knowledge (Grafana, Opera, Splunk)
- Linux fundamentals (networking, CPU, memory, storage)
- strong communication
- collaboration
- problem-solving
- Hadoop
- Kafka
- AWS EMR