At IBM Software, we transform client challenges into solutions. Building the world’s leading AI-powered, cloud-native products that shape the future of business and society. Our legacy of innovation creates endless opportunities for IBMers to learn, grow, and make an impact on a global scale. Working in Software means joining a team fueled by curiosity and collaboration. You’ll work with diverse technologies, partners, and industries to design, develop, and deliver solutions that power digital transformation. With a culture that values innovation, growth, and continuous learning, IBM Software places you at the heart of IBM’s product and technology landscape. Here, you’ll have the tools and opportunities to advance your career while creating software that changes the world. With Confluent, data doesn’t sit still. We put information in motion, streaming in near real time so organizations can react faster, build smarter, and deliver experiences as dynamic as the world around them.
The Secure Compute Team builds and operates the foundational infrastructure layer that powers Confluent Cloud. Our mission is to enable the secure execution of code and data processing in multi-tenant environments across AWS, Azure, and GCP at scale. We provide the essential building blocks, including isolation, identity, and networking, that empower both our customers and internal product teams to innovate without compromising security.
As Confluent's product portfolio continues to expand, our team sits at the center of the company's growth strategy. We own a polyglot, container-based runtime that operates across thousands of clusters globally, solving some of the most complex challenges in distributed systems and cloud-native architecture. By balancing world-class security with operational efficiency, we ensure Confluent Cloud remains the trusted platform for mission-critical workloads, even in the most highly regulated industries.
About the RoleAs a Staff Software Engineer on the Secure Compute Platform team, you will be a key technical leader responsible for building and evolving a next-generation, multi-tenant, cloud-native compute platform that safely runs both trusted and untrusted workloads at scale. Our platform is built on Kubernetes and operates across a large fleet of clusters spanning multiple public clouds, providing a unified abstraction layer for workload execution, lifecycle management, security, and operational excellence.
What You Will DoDefine and drive the technical direction for Secure Compute, including platform architecture, runtime capabilities, and security models for running trusted and untrusted workloads at scale
Design and implement platform APIs and Kubernetes controllers/operators, primarily in Go, that power workload lifecycle management, autoscaling, placement, and workload isolation for containers and serverless-style functions
Partner with product and platform teams to shape and deliver the Secure Compute roadmap, enabling new customer-facing capabilities and internal platforms to build on a common compute substrate
Lead high-impact initiatives in areas such as workload scheduling, failure and disruption handling, private and public networking architectures, deployment and rollout strategies, and fleet-wide resource management
Drive technical design reviews and influence architectural decisions across teams, ensuring Secure Compute services are easy to adopt, secure by default, and aligned with broader platform objectives
Mentor and develop engineers through technical leadership, design guidance, code reviews, pair programming, and the promotion of best practices for secure, reliable, and scalable platform development
Own operational excellence for critical Secure Compute services, including availability, reliability, performance, service level objectives (SLOs), on-call operations, incident response, and disaster recovery
**This role can be performed remotely from anywhere within the United States.
10+ years of relevant experience delivering scalable backend, infrastructure or platform software in production
Bachelor's degree in Computer Science or a related field, or equivalent practical experience
Proven experience building and operating large-scale, highly available distributed systems
Self-starter with strong problem-solving skills and the ability to thrive in a fast-paced environment
Deep expertise in Kubernetes, including controller development, operator patterns, and preferably multi-region and multi-cluster architectures
Strong proficiency in Go, Scala, C++, or other statically typed programming languages, with experience building production-grade services and control planes.
Experience with multi-tenant platform architectures and security/isolation patterns (for example, namespaces, network policies, sandboxing, secrets management, identity and access management)
Hands-on experience with secure container runtimes and Linux internals (for example, Kata containers, Cloud Hypervisor, cgroups, namespaces, seccomp)
Experience troubleshooting and optimizing performance for containerized and virtualized workloads
Familiarity with gRPC, Protocol Buffers (Protobuf), and platform API design for service-to-service communication
Experience working with public cloud platforms and integrations, including AWS, Google Cloud Platform (GCP), and Microsoft Azure
Strong collaboration skills with a track record of partnering effectively across Product, SRE/Operations, Security, and Engineering teams
Demonstrated technical leadership and mentorship experience, including driving cross-functional alignment on architecture, strategy, and execution
Experience in one or more of the following domains: storage systems, computer orchestration, networking, security engineering, performance engineering
Familiarity with Kubernetes ecosystem technologies, including service meshes and modern cloud-native architectures
Contributions to open-source infrastructure or cloud platform projects