Overview In this role you will contribute to research, design and development of large-scale foundation models for machine-generated data. You will focus on graph data while supporting logs, time series, traces, and events, collaborating with cross-functional teams to meet business needs. Your work will aim to improve model quality, scalability and operational efficiency in real-time enterprise environments. This is an opportunity to shape AI foundation models that enhance reliability, security and predictive insights at scale. You will join a culture that values technical excellence and impact-driven collaboration, with the chance to influence Splunk Observability, Security and Platform initiatives.
Compensation / Benefits- medical, dental and vision insurance
- 401(k) with company match
- paid parental leave
- paid time off and flexible vacation
- restricted stock units (RSUs)
- volunteer days
Responsibilities- Research and develop large-scale foundation models for machine-generated data, emphasizing graph data
- Design and enhance distributed training and inference workflows to improve scalability and efficiency
- Collaborate with engineering, product and data science teams to translate requirements into AI/ML solutions
- Share technical insights and best practices to advance team capabilities and project outcomes
- Explore and evaluate new AI/ML techniques and tools to support roadmaps and objectives
- Own projects end-to-end, identify obstacles and improve development processes with urgency
Key requirements- Master in Computer Science or related quantitative field + 2+ years of industry research
- Proven track record in graph representation learning or large-language modeling or multi-modal fusion
- Strong Python skills and proficiency with deep learning frameworks (PyTorch, TensorFlow)
- Experience translating research ideas into production systems
- Collaborative communication
- Proactive problem solving
- Ownership and accountability
- Graph neural networks (GCN, GAT, GraphSAGE) and related graph transformers
- Large language modeling for structured and unstructured data
- Multi-modal fusion of graph, text, log, and time-series data