Updating your browser will give you an optimal website experience. Learn more about our supported browsers.
IT Infrastructure Lead
What we need
TCDRS offers an innovative, strategic and collaborative culture that uses best-practice technologies to enable our strategic initiatives, services, security and business processes. We have been focused on delivering services in a more digital, efficient and secure environment. Over the past few years, we’ve worked to modernize our systems and processes — not just to keep up with technology, but to better meet the evolving needs of our members and employers.
One of our most significant shifts has been implementing automated workflows and straight-through processing when coded validations pass to improve security, efficiency, flexibility and empowerment. Read more about our digital transformation story.
The Infrastructure Lead is a hands-on technical leader responsible for the design, implementation, operation and continuous improvement of the organization’s core IT infrastructure. This role leads a small team of systems administrators while remaining actively engaged in operational support and infrastructure modernization initiatives.
The Infrastructure Lead ensures systems are secure, reliable, scalable, recoverable and aligned with business and security requirements. The position collaborates across infrastructure, security, service desk, web and application development teams to deliver resilient, high-performing platforms that support business operations.
What you’ll do
1. Team Leadership and Development
-
Lead, coach and support a small team of systems administrators, establishing clear priorities, expectations and accountability.
-
Assign and balance operational work, projects and escalations based on business priorities and team capacity.
-
Conduct regular team meetings, provide timely feedback and support professional development and cross-training.
-
Provide technical direction and decision support during incidents, projects, architecture reviews and infrastructure changes.
-
Communicate team progress, risks, dependencies and operational concerns clearly to IT leadership.
-
Promote a collaborative, security-first culture focused on ownership and continuous improvement.
2. Infrastructure Design and Operations
-
Design, implement and maintain enterprise infrastructure, including servers, storage, virtualization, backup, identity and other core services.
-
Demonstrate a strong understanding of enterprise network infrastructure, including switches, routers, firewalls and Wi-Fi technologies, and how they integrate with core systems and services.
-
Administer Windows and Linux environments across on-premises, cloud and hybrid architectures.
-
Monitor system performance, availability, capacity and health; proactively identify and resolve issues.
-
Plan and execute system upgrades, patching, lifecycle management and technology refreshes.
-
Drive infrastructure standardization, modernization, configuration consistency and technical debt reduction.
3. Cross-Team Collaboration and Enablement
-
Provide technical consultation and infrastructure support to web and application teams to ensure platform and hosting requirements are met.
-
Collaborate on system architecture, deployments, performance tuning, availability planning and production readiness.
-
Act as the infrastructure subject matter expert during application design reviews and major releases.
-
Troubleshoot complex, cross-stack issues spanning infrastructure, operating systems, middleware, networks and applications.
-
Partner with stakeholders to meet performance, security and recoverability requirements.
-
Apply strong requirements-gathering, technical evaluation, vendor coordination, procurement support, implementation planning and communication skills when working with vendors to acquire and implement new IT products and services.
4. Security and Compliance
-
Implement and manage secure key, secret and certificate lifecycle processes, including generation, storage, rotation, renewal and revocation.
-
Administer and support enterprise key management solutions such as Azure Key Vault or equivalent platforms.
-
Protect encryption keys, application secrets, API keys, service credentials and certificates used across infrastructure, applications and cloud services.
-
Monitor certificate expiration and automate renewal processes to reduce the risk of service disruption.
-
Partner with security, application and web teams to integrate secure configuration and secret management practices into system and application design.
-
Support vulnerability remediation, access reviews, audit requests, policy compliance and risk-reduction initiatives.
5. Operational Support, Incident Management, and Documentation
-
Provide Tier 3 escalation support and lead root cause analysis for complex or recurring system issues.
-
Coordinate technical response and recovery activities during significant infrastructure incidents.
-
Collaborate with network, security, application and service desk teams to resolve incidents and improve service reliability.
-
Develop and maintain technical documentation, architecture diagrams, standards and procedures.
-
Identify opportunities to improve operational efficiency, reliability and consistency through automation and process improvement.
-
Use operational metrics, recurring issue trends and lessons learned to prioritize service improvements.
6. Disaster Recovery and Business Continuity
-
Design, implement and maintain disaster recovery strategies for critical systems and applications.
-
Ensure backup, replication and recovery solutions align with defined recovery time objective (RTO) and recovery point objective (RPO) requirements.
-
Develop and maintain disaster recovery documentation, runbooks, system dependency mappings and recovery procedures.
-
Plan and participate in disaster recovery testing, tabletop exercises and failover and failback events.
-
Partner with application, web, security and business stakeholders to align recovery capabilities with business requirements.
-
Identify, communicate, and remediate gaps in redundancy, backup coverage and recoverability.
What you should have
Required Qualifications
-
Bachelor’s degree in information technology, Computer Science or a related field, or equivalent practical experience.
-
Five or more years of progressive experience in systems administration, infrastructure operations or a related technical role.
-
Experience leading technical work, mentoring systems administrators and coordinating team priorities.
-
Strong experience with Windows Server and Workstation operation systems.
-
Experience with virtualization technologies such as VMware, Hyper-V or equivalent platforms.
-
Experience with identity and access management technologies, including Active Directory and Microsoft Entra ID, or equivalent platforms.
-
Experience supporting on-premises, cloud or hybrid environments.
-
Working knowledge of networking fundamentals, including DNS, DHCP, TCP/IP, routing concepts and firewalls.
-
Experience with backup, recovery, system monitoring and infrastructure management solutions.
-
Experience with key, secret and certificate management solutions such as Azure Key Vault or similar platforms.
-
Understanding of public key infrastructure (PKI), certificate authorities and certificate lifecycle management.
Preferred Qualifications
-
Experience with Microsoft Azure and Microsoft 365 environments.
-
Experience with security tooling, including endpoint detection and response and vulnerability management platforms.
-
Automation and scripting experience using PowerShell or similar tool.
-
Experience supporting application hosting environments, middleware or web platforms.
-
Relevant Microsoft, VMware, cloud, IT service management or security certifications.
Key Competencies
-
Hands-on leadership with the ability to guide a team while contributing directly to technical work.
-
Demonstrated ability to set priorities, delegate effectively and hold team members accountable for outcomes.
-
Ability to develop people through coaching, feedback, knowledge sharing and growth opportunities.
-
Strong ownership of initiatives from concept through completion with limited oversight.
-
Ability to manage competing operational and project priorities in a dynamic environment.
-
Strong cross-functional collaboration, stakeholder communication and conflict-resolution skills.
-
Calm, structured decision-making during incidents and other high-pressure situations.
-
Strong attention to detail, documentation, organization and operational follow-through.
-
Security-first approach to infrastructure design and operations.
What you’ll get
TCDRS offers a competitive salary, exceptional health benefits, and a collaborative work environment in Austin, Texas, overlooking Zilker Park. Our hybrid schedule includes four days in the office and one remote day each week. At this time, we are unable to provide visa sponsorship for this position.
At TCDRS, you will have the opportunity to lead meaningful technology initiatives that directly improve how hundreds of thousands of Texans access and manage important retirement benefits while helping shape the future of a highly innovative technology organization.
We help provide more than 400,000 Texans with retirement, disability and survivor benefits. TCDRS is one of the best-funded retirement systems in the nation. Although we provide retirement benefits or “pensions” to hard-working Texans, our unique features distinguish us from traditional pension plans and keep us financially strong.
Our culture is grounded in integrity, care and anticipation. We believe in doing the right thing, serving our members and employers with care, and looking ahead to create better solutions for the future.
Apply Now
Submit your resume to employment@tcdrs.org.
Please refer to the job title in the subject of your email.