Senior Dev Ops Engineer
University of Chicago · Chicago, IL
📍 Chicago, ILvia workdayFirst listed here 2026-09-24
Apply on company site ↗
Career Moonshot pulls this listing straight from the employer's hiring system — no recruiter middleman, no reposts. Applying takes you directly to University of Chicago.
Department
Globus Systems Operations
About the Department
Globus (www.globus.org) is a sustainable, non-profit unit within The University of Chicago delivering solutions to the research community worldwide. Globus develops and provides critical services that support scientific research for governmental, academic, and commercial organizations in a wide range of disciplines including life sciences, physics, and astronomy. We develop and operate commercial-quality, cloud-based software application and platform services used by 10s of thousands of researchers to manage their large–and growing–data management challenges. We have offices located at the NBC Tower in the heart of downtown Chicago and remote employees who work-from-home. Globus, together with Globus Labs, a research group within the University of Chicago, and part of the Data Science and Learning Division at Argonne National Labs, develop and deploy cutting edge technologies to solve new challenges facing the scientific community and enable break-through scientific discoveries.
Job Summary
The Globus Operations team is a 4-5 person group that designs, builds, and operates the cloud infrastructure behind a research-computing platform used by hundreds of institutions worldwide. Globus is a hybrid solution combining AWS-hosted orchestration services with installable applications, delivering identity and access management, data transfer and sharing, and task automation as both software-as-a-service (SaaS) and platform-as-a-service (PaaS).
As a senior member of this small, high-autonomy team, you will ensure software development and operational best practices are effectively integrated across our services and AWS infrastructure. The team owns an unusually broad estate — an AWS Organization of roughly thirty production accounts, a golden-image pipeline producing around thirty machine images, a self-hosted monitoring platform, a centralized security-logging pipeline, and in-house compliance and security automation — all managed as code in Python, Bash, Terraform, CloudFormation, and Ansible. You will work both collaboratively and independently on complex issues and projects, in some cases making progress under minimal guidance. You will also embed with select software engineering teams to provide operational and infrastructure guidance on best practices.
We are looking for a senior engineer equally comfortable improving a build pipeline and debugging a production incident, and who will be a leader and educator both within and across teams.
Responsibilities
Architecture and Design: Participate in the definition and documentation of cloud infrastructure architecture — networking, monitoring, logging, security, backup, and deployment — across production and development environments. Design new systems and tools, identify opportunities for technical improvement, and review and test solutions to ensure standards are met.
Software Development: Develop, test, document, and maintain high-quality software and infrastructure as code for deploying and operating Globus cloud services, in Python, Bash, Terraform, and CloudFormation. Contribute to shared internal libraries and expand automated test coverage of operational tooling.
Build and Release Automation: Maintain and evolve the machine-image and deployment supply chain — Packer-built AMIs produced through AWS CodeBuild, published via SSM Parameter Store, KMS-encrypted and shared cross-account.
Monitoring and Observability: Maintain the current monitoring and alerting systems (self-hosted Nagios/NRPE with configuration generated from live AWS inventory, CloudWatch alarms, PagerDuty routing) and help in the migration to Prometheus.
Security Operations and Compliance: Contribute to Globus's security plan, controls, monitoring, and reporting. Extend in-house continuous-compliance automation, operate intrusion detection and endpoint protection tooling, maintain hardened machine images.
Identity, Access, and Trust: Administer IAM roles, policies, and cross-account trust across the AWS Organization.
SRE/Operations: Deploy, operate, and monitor production Globus services for high availability. Participate in an on-call rotation, lead incident response, and author root-cause analyses. Identify recurring operational toil and eliminate it through automation and better defaults.
Support: Act as a technical consultant and resource for other team members, including the engineering and user-support teams, assisting with operational issues and troubleshooting.
Designs new systems, features, and tools. Solve complex problems and identify opportunities for technical improvement and performance optimization. Review and test code to ensure appropriate standards are met.
Utilize technical knowledge of existing and emerging technologies, including public cloud offerings from Amazon Web Services, Microsoft Azure, and Google Cloud.
Performs other related work as needed.
Minimum Qualifications
Education:
Minimum requirements include a college or university degree in related field.
Work Experience:
Minimum requirements include knowledge and skills developed through 5-7 years of work experience in a related job discipline.
Certifications:
---
Preferred Qualifications
Experience:
Strong experience operating AWS at scale — IAM, VPC (peering, endpoints, security groups), EC2 and Auto Scaling Groups, ECS (EC2 and Fargate), S3, RDS, Route 53, Lambda, Systems Manager, CloudWatch — including cross-account work in a multi-account AWS Organization.
Strong experience building Infrastructure as Code with Terraform (modules, remote state, provider upgrades, reconciling drift against live environments), the ability to maintain and incrementally retire legacy CloudFormation is a plus.
Experience automating infrastructure in Python — maintainable, reusable libraries rather than only scripts, using boto3 — and in Bash.
Building Amazon machine images and the instances they run on, plus Docke
More Chicago, IL jobs
Chicago, IL jobs · Browse all locations