We use cookies. Find out more about it here. By continuing to browse this site you are agreeing to our use of cookies.
#alert
Back to search results
New

Research Computing Engineer (AI Infrastructure and HPC)

The Pennsylvania State University
paid time off, sick time, remote work
United States, Pennsylvania, University Park
201 Old Main (Show on map)
Sep 12, 2026
APPLICATION INSTRUCTIONS:
  • CURRENT PENN STATE EMPLOYEE (faculty, staff, technical service, or student), please login to Workday to complete the internal application process. Please do not apply here, apply internally through Workday.
  • CURRENT PENN STATE STUDENT (not employed previously at the university) and seeking employment with Penn State, please login to Workday to complete the student application process. Please do not apply here, apply internally through Workday.
  • If you are NOT a current employee or student, please click "Apply" and complete the application process for external applicants.

Approval of remote and hybrid work is not guaranteed regardless of work location.For additional information on remote work at Penn State, seeNotice to Out of State Applicants.

POSITION SPECIFICS

The Institute for Computational and Data Sciences (ICDS) at Penn State seeks a Research Computing Engineer to join our technical team. This role supports Penn State's research mission by designing, operating, automating, and optimizing the GPU and AI and computing infrastructure used by researchers across the university, along with the high-performance computing systems that support it.

The position will be filled at the Reserach Computing Systems Engineer - Advanced Professional level.

Candidates must be U.S. citizens due to specific access requirements associated with this position.

This position is ideal for an engineer who enjoys building reliable, scalable systems for machine learning, data-intensive research, and advanced computing workloads. The successful candidate will work across GPU systems, HPC platforms, storage, networking, automation, and user-facing research workflows to enable cutting-edge research in AI, simulation, and computational science.

Work Arrangement:This is a full-time position, which will report to the Research HPC Manager and requires on-site work at University Park and is not supportive of remote work.

Responsibilities: As part of a collaborative engineering team, you'll contribute across a broad range of responsibilities that include the following:

  • Collaborate with teammates, users, and vendor support to diagnose issues and implement solutions across compute, storage, networking, and software environments.

  • Monitor, maintain, automate, and improve AI and HPC systems and supporting infrastructure.

  • Design, deploy, operate, troubleshoot, and optimize systems using DevOps and infrastructure-as-code practices.

  • Support GPU-accelerated computing environments for AI, machine learning, and scientific workloads.

  • Partner with researchers and ICDS staff to understand workload requirements and develop practical engineering solutions for system configuration, performance, and research workflows.

  • Support security, logging, documentation, and compliance process for the systems we operate, including environments subject to federal research security.

  • Contribute to planning, requirements gathering, process improvement, and operational readiness for new services and infrastructure.

  • Provide timely updates to system documentation and respond to user questions with clear, actionable guidance.

  • Evaluate and improve tools, platforms, and workflows that support AI model development, training, inference, and data movement at scale.

Required qualifications and skills include the following:

  • Administration of multi-GPU nodes at scale, including driver and firmware lifecycle management, NVLink/NVSwitch topology validation, GPU health monitoring and tuning for multi-GPU or multi-node GPU workloads.

  • Ability to work effectively in a Linux environment, including command-line tools, file editing, POSIX permissions, and system configuration.

  • Strong scripting ability in Bash and Python.

  • Strong problem-solving skills and the ability to debug complex systems.Ability to work collaboratively as part of a technical team.Experience using AI tools or AI agents to improve programming, debugging, development, or prototyping workflows.

  • Clear written and verbal communication skills.

Preferred qualifications:

Experience with one or more of the following is helpful but not required:

AI and HPC Workloads
  • Experience supporting AI/ML infrastructure, including environments used for model training, inference, experiment workflows, and large-scale data processing.

  • Experience with job schedulers such as Slurm, PBS, HTCondor, or LSF.

  • Software development experience and familiarity with HPC programming environments such as C/C++, Fortran, CUDA, MPI, or OpenMP.

Infrastructure and Automation
  • DevOps experience, including Git-based workflows, CI/CD, automation, and collaborative development practices.

  • Experience with system deployment tools such as xCAT, Warewulf, OpenCHAMI, OpenStack/Bifrost, or MAAS.

  • Experience administering or supporting Kubernetes.

  • Experience with virtualization and containerization technologies such as VMware, Docker, Apptainer, or Podman.

Networking and Storage
  • Networking experience, including EVPN, BGP, and IPv6.

  • Experience with high-speed interconnects such as InfiniBand, HPE Slingshot, or similar HPC fabrics.

  • Experience with HPC or distributed storage systems such as GPFS, Lustre, Ceph, or VAST.

Observability and Security
  • Monitoring and observability experience with tools such as Grafana, Graphite, Prometheus, VictoriaMetrics, or InfluxDB.

  • Experience with databases such as MySQL/MariaDB or PostgreSQL.

  • Security experience including identity and access management, single sign-on, and LDAP/Active Directory.

Background and Process
  • Familiarity with Agile project development.

  • Prior experience in academic research computing, research data infrastructure, or large-scale shared computing environments.

Application Instructions: To receive full consideration for the position, applications should include:

  • A cover letter expressing the candidate's interest in the role
  • A current Curriculum Vitae (CV) or Resume

Why ICDS

At ICDS, you'll build and support infrastructure that enables advanced research in artificial intelligence, machine learning, simulation, data science, and computationally intensive discovery. You'll work on clusters, storage and GPU systems Penn State researchers rely on for AI, machine learning and computational research. You won't own a narrow slice of that. Our engineers work across provisioning, scheduling, storage, networking, GPU infrastructure and user-facing tooling. You'll debug problems from a user's job submission all the way down to drivers and system firmware, and you'll learn a lot on the way.

You'll also be close to the science. You'll talk to the people running the workloads, understand what they're trying to accomplish, and engineer systems around actual research needs rather than requirements handed down from somewhere else. It's a collaborative team, and the problems are rarely the same twice.

    MINIMUM EDUCATION, WORK EXPERIENCE & REQUIRED CERTIFICATIONS

    Bachelor's Degree 3+ years of relevant experience; or an equivalent combination of education and experience accepted Required Certifications: None

    BACKGROUND CHECKS/CLEARANCES

    Employment with the University will require successful completion of background check(s) in accordance with University policies. Penn State does not sponsor or take over sponsorship of a staff employment Visa. Applicants must be authorized to work in the U.S.

    SALARY & BENEFITS

    The salary range for this position, including all possible grades, is $91,488.00 - $137,280.00.

    Salary Structure - Information on Penn State's salary structure

    Penn State provides a competitive benefits package for full-time employees designed to support both personal and professional well-being. In addition to comprehensive medical, dental, and vision coverage, employees enjoy robust retirement plans and substantial paid time off which includes holidays, vacation and sick time. One of the standout benefits is the generous 75% tuition discount, available to employees as well as eligible spouses and children. For more detailed information, please visit our Benefits Page.

    CAMPUS SECURITY CRIME STATISTICS

    Pursuant to the Jeanne Clery Disclosure of Campus Security Policy and Campus Crime Statistics Act and the Pennsylvania Act of 1988, Penn State publishes a combined Annual Security and Annual Fire Safety Report (ASR). The ASR includes crime statistics and institutional policies concerning campus security, such as those concerning alcohol and drug use, crime prevention, the reporting of crimes, sexual assault, and other matters. The ASR is available for review here.

    EEO IS THE LAW

    Penn State is an equal opportunity employer and is committed to providing employment opportunities to all qualified applicants without regard to race, color, religion, age, sex, sexual orientation, gender identity, national origin, disability or protected veteran status. If you are unable to use our online application process due to an impairment or disability, please contact 814-865-1473.

    Penn State is committed to and accountable for advancing equity, respect, and belonging. We embrace individual uniqueness, as well as a culture of belonging that supports equity initiatives, leverages the educational and institutional benefits of inclusion in society, and provides opportunities for engagement intended to help all members of the community thrive. We value belonging as a core strength and an essential element of the university's teaching, research, and service mission.

    Federal Contractors Labor Law Poster

    PA State Labor Law Poster

    Penn State Policies

    Copyright Information

    Hotlines

    Applied = 0

    (web-665cd84569-4dpxh)