Systems Administrator Mid-Level Interview Preparation Guide (FAANG Standards)
This guide is based on general FAANG interview practices and may not reflect specific company procedures.
FAANG companies conduct comprehensive interview processes for mid-level Systems Administrator roles consisting of an initial recruiter screen, technical assessments across multiple infrastructure domains, hands-on troubleshooting scenarios, infrastructure design discussions, and behavioral evaluations. The process assesses depth of systems knowledge, practical troubleshooting abilities, architectural thinking, and alignment with company leadership principles.
Interview Rounds
Recruiter Screening
What to Expect
Initial conversation with recruiter to assess background, motivation, and cultural fit. This is a 30-45 minute call to verify your interest in the role, confirm your experience level matches mid-level expectations, and assess communication skills and professionalism. The recruiter will review your resume and probe into your infrastructure background, project scope, and why you're interested in the Systems Administrator role at their company.
Tips & Advice
Have a clear narrative about your infrastructure journey and why you're at a mid-level. Highlight 2-3 projects where you made meaningful contributions beyond day-to-day operations. Be specific about infrastructure domains you specialize in (e.g., Linux systems, virtualization, backup systems). Show enthusiasm for the company and mention specific aspects of their tech culture or infrastructure challenges that appeal to you. Practice concise answers—recruiters typically have limited time. Prepare questions about the team structure, infrastructure stack, and career development opportunities.
Focus Topics
Communication and Professional Presence
Demonstrate clarity, professionalism, and thoughtfulness in communication. Speak at a measured pace. Avoid filler words ('um', 'like'). Answer questions directly but with appropriate depth. Ask intelligent follow-up questions. Show genuine interest through tone and engagement.
Practice Interview
Study Questions
Motivation and Company Fit
Articulate genuine reasons for interest in the role and company. Research the company's infrastructure challenges, scale, and technology stack. Show understanding of how their business depends on reliable infrastructure. Explain what attracts you to their specific environment and how this role aligns with your career goals.
Practice Interview
Study Questions
Career Progression and Experience Narrative
Develop a clear story of your progression from entry-level to mid-level systems administrator. Identify key projects and responsibilities that demonstrate growing ownership, mentoring of junior staff, and contribution to infrastructure decisions. Be prepared to explain what makes you mid-level versus junior or senior.
Practice Interview
Study Questions
Infrastructure Domain Expertise Areas
Clearly articulate the specific infrastructure domains where you have depth: Linux systems, Windows servers, virtualization, networking, storage, monitoring, backup/disaster recovery, cloud platforms, or security. Quantify your experience (e.g., 'managed 150 Linux servers', 'designed backup strategy for 50+ applications'). Be honest about gaps.
Practice Interview
Study Questions
Technical Phone Screen
What to Expect
60-minute technical assessment with a senior systems administrator or infrastructure engineer. This round tests foundational knowledge across operating systems, networking, and basic infrastructure concepts. You'll answer direct technical questions, discuss troubleshooting approaches, and demonstrate systems thinking. The interviewer assesses breadth of knowledge, ability to explain concepts clearly, and problem-solving methodology. This is not a hands-on coding round but may include command-line scenarios or architecture sketches.
Tips & Advice
Review operating system fundamentals: processes vs threads, memory management (paging, segmentation, virtual memory), CPU scheduling basics, and concurrency concepts. Study networking basics: TCP/IP stack, DNS, HTTP/HTTPS, common ports, routing concepts. Be ready to discuss Linux and Windows differences for common administrative tasks. When answering questions, explain your thinking process, not just the answer. For scenario questions (e.g., 'How would you troubleshoot high CPU usage?'), walk through a systematic approach: check metrics, identify problematic process, investigate root cause, implement fix, verify. Ask clarifying questions if scenarios are ambiguous. Prepare examples from your past where you diagnosed and resolved infrastructure problems—use the STAR method (Situation, Task, Action, Result). If you don't know an answer, say so honestly and discuss how you would learn it.
Focus Topics
Basic Linux and Windows Command Line
Essential commands for system administration on both platforms. Linux: ls, cd, mkdir, rm, chmod, chown, grep, find, sed, awk, tail, head, ps, top, netstat, systemctl, journalctl, sudo, ssh. Windows: dir, cd, mkdir, del, icacls, Get-Process, Get-Service, Get-EventLog, powershell basics, net commands. Be comfortable using these interactively.
Practice Interview
Study Questions
System Monitoring and Performance Analysis
Understanding key metrics: CPU utilization, memory consumption, disk I/O, network bandwidth, process monitoring. Tools: top/htop, ps, vmstat on Linux; Task Manager, Performance Monitor, Get-Process on Windows. Interpret metrics to diagnose problems. Understand load averages on Linux, context switch rates, page faults, and I/O wait time.
Practice Interview
Study Questions
Operating System Fundamentals (Linux and Windows)
Core concepts in both Linux and Windows: process and thread management, memory management (paging, virtual memory, swapping), CPU scheduling, interrupt handling, system calls, kernel vs user space. Understand how these concepts differ between Linux and Windows. Be able to explain deadlock conditions, context switching, and resource contention.
Practice Interview
Study Questions
Networking Fundamentals
TCP/IP stack layers, DNS resolution process, HTTP/HTTPS, common ports and protocols (SSH, FTP, SMTP, NTP), routing concepts, network interfaces, IP addressing (IPv4 and IPv6), and firewall basics. Understand how to diagnose network connectivity issues using tools like ping, traceroute, netstat, and tcpdump on Linux; ipconfig, netstat, tracert on Windows.
Practice Interview
Study Questions
Troubleshooting Methodology and Systematic Problem-Solving
Develop a systematic approach to troubleshooting: gather information, establish a baseline, generate hypotheses, test hypotheses methodically, implement fix, verify resolution. Practice articulating this approach clearly. Discuss how you would diagnose common infrastructure issues: high CPU usage, memory leaks, disk space issues, network connectivity problems, service failures, performance degradation.
Practice Interview
Study Questions
Linux Systems Administration Deep Dive
What to Expect
90-minute technical interview focusing specifically on Linux systems administration. You'll be assessed on practical Linux knowledge, command-line proficiency, user and permission management, package management, service/daemon management, shell scripting basics, and system administration patterns. Expect detailed questions about how you would configure and maintain Linux systems in a production environment. May include live terminal scenarios where you demonstrate commands or explain shell scripts.
Tips & Advice
Practice using the Linux command line on real systems or virtual machines. Build practical experience with user management (useradd, userdel, usermod, passwd), file permissions (chmod, chown, umask), and sudoers configuration. Understand package managers (apt/apt-get on Debian, yum on Red Hat) and how to troubleshoot dependency issues. Study systemd and systemV init systems, including how to write basic service files, enable/disable services, and check service status. Write simple bash scripts for common administration tasks: log rotation, user creation, backup verification, system health checks. Be ready to read and modify existing shell scripts. Understand important Linux directories (/etc, /var, /opt, /home, /root, /proc, /sys) and what they contain. Review Linux file systems and permissions model deeply. For any code or configuration shown, be prepared to explain what it does and identify potential issues. If asked to write a script or configuration, think aloud about requirements and edge cases.
Focus Topics
Linux Logging, Log Analysis, and System Journals
Understanding syslog architecture and rsyslog configuration. Log file locations (/var/log). Using journalctl for systemd logs. Log rotation with logrotate. Parsing and analyzing logs: grep, tail, less, awk. Forwarding logs to centralized systems. Understanding common log locations: auth, syslog, kernel messages, application logs. Troubleshooting using logs.
Practice Interview
Study Questions
Linux Bash Shell and Shell Scripting
Bash basics: variables, arrays, conditionals (if/else), loops (for, while), functions. String manipulation and regular expressions. Input/output redirection and pipes. Common utilities: grep, sed, awk, cut, sort, uniq. Error handling and exit codes. Writing scripts for common administration tasks: system health checks, automated backups, log rotation, user provisioning, monitoring. Best practices: error handling, logging, input validation, documentation.
Practice Interview
Study Questions
Linux Networking Configuration and Troubleshooting
Network interfaces and configuration files. IP addressing and routing configuration. Using netstat, ss, ip, ifconfig commands. DNS configuration and troubleshooting. Network connectivity testing (ping, traceroute, telnet, curl). Network statistics and monitoring. Firewall basics (iptables/firewalld). Understanding network namespaces.
Practice Interview
Study Questions
Linux Package Management and System Updates
Package managers: apt/apt-get on Debian/Ubuntu, yum/dnf on Red Hat/CentOS. Installing, updating, removing packages. Understanding dependencies. Repository management. Security updates vs regular updates. Kernel updates and testing. Patching strategies for production systems. Rolling updates and minimizing downtime. Version pinning and compatibility management.
Practice Interview
Study Questions
Linux System Services and Process Management
Understanding systemd (modern systems) and systemV init (legacy). Creating and modifying systemd unit files. Service lifecycle: start, stop, enable, disable, restart, reload. Process monitoring: ps, top, htop, pgrep, pkill. Understanding process states (running, sleeping, stopped, defunct). Dealing with zombie processes. Managing background jobs, signals (SIGTERM, SIGKILL), and graceful shutdowns.
Practice Interview
Study Questions
Linux File Permissions and Access Control
Understanding Linux permission model: rwx permissions, octal notation, special bits (setuid, setgid, sticky bit). Using chmod, chown, chgrp. Understanding umask and default permissions. File Access Control Lists (ACLs) when appropriate. Understanding ownership vs permissions. Practical scenarios: securing sensitive files, allowing group access, preventing accidental deletion.
Practice Interview
Study Questions
User and Group Management on Linux
Creating, modifying, and deleting users and groups. Understanding /etc/passwd, /etc/shadow, /etc/group, /etc/gshadow files. Using useradd, userdel, usermod, groupadd, groupdel commands. Configuring default shell, home directory, UID/GID management. Understanding and configuring sudoers file. Best practices for user account lifecycle management in enterprise environments. Password policies and aging.
Practice Interview
Study Questions
Windows Server Administration Deep Dive
What to Expect
90-minute technical interview focusing on Windows Server administration. You'll be assessed on Windows Server architecture, Active Directory and Group Policy, PowerShell scripting, Windows services management, server roles and features, and Windows security. Expect detailed questions about managing Windows infrastructure at scale. May include scenarios where you design or troubleshoot Windows environments, or write PowerShell scripts for common administration tasks.
Tips & Advice
Deep dive into Active Directory: forest/domain structure, organizational units, group policies, group policy objects (GPOs), domain controllers, replication. Understand Group Policy processing and troubleshooting common GPO issues. Get comfortable with Active Directory Users and Computers management console and PowerShell AD cmdlets. Study Windows Server roles and features: file services, print services, remote access, domain controllers. Understand PowerShell as a core Windows administration tool—practice writing scripts for user/group management, system configuration, monitoring, and reporting. Learn key Windows utilities: Event Viewer for troubleshooting, Performance Monitor for performance analysis, Services snap-in for service management. Understand Windows file sharing (SMB), NTFS permissions, and share permissions. Study Windows security: user account control (UAC), Windows Defender, Windows Firewall, patch management through WSUS or update management tools. Be ready to troubleshoot common Windows server issues: service failures, Group Policy problems, user login issues, performance problems, disk space issues.
Focus Topics
Windows Performance Monitoring and Troubleshooting
Performance Monitor: creating data collectors, analyzing performance counters. Task Manager for quick assessment. Event Viewer for troubleshooting. Resource Monitor. Common performance issues: high CPU, high memory consumption, disk bottlenecks, network issues. Baseline performance and detecting anomalies. Troubleshooting slow performance, application crashes, and service failures.
Practice Interview
Study Questions
Windows Server Roles, Features, and Services
Core Windows Server roles: Domain Controller, File Server, Print Server, Remote Access, DNS, DHCP. Installing and configuring roles using Server Manager or PowerShell. Understanding role services and dependencies. Managing services: automatic startup, manual, disabled. Service dependencies and startup order. Troubleshooting service failures using Event Viewer and Services console.
Practice Interview
Study Questions
Windows File Sharing, NTFS Permissions, and Access Control
SMB protocol basics. Creating and managing file shares. Understanding NTFS permissions vs share permissions. Permission inheritance and effective permissions. Setting permissions on files and folders. Troubleshooting access denied errors. Taking ownership of files. Using icacls command-line tool. Understanding special permissions.
Practice Interview
Study Questions
Windows Security, Patching, and Compliance
Windows security fundamentals: authentication, authorization, auditing. Windows Defender and antimalware configuration. Windows Firewall configuration and rules. Windows Update and WSUS for patch management. Security baselines and hardening. Audit policy configuration. User Account Control (UAC). Credential Guard and other modern security features. Security Event log monitoring.
Practice Interview
Study Questions
Group Policy and System Configuration Management
Group Policy Objects (GPOs): creation, linking, inheritance, and filtering. Group Policy scope: site, domain, OU. User vs computer policies. Common policy settings: password policy, audit policy, security settings, registry settings, script deployment. Group Policy processing and order of application. Troubleshooting GPO issues: gpresult, gpupdate, Event Viewer. Security Group Policy: user rights assignments, audit policies, account policies. Using Group Policy to manage Windows Update, firewall, disk encryption.
Practice Interview
Study Questions
PowerShell for Windows Server Administration
PowerShell fundamentals: cmdlets, pipeline, objects, and properties. Common AD cmdlets: Get-ADUser, Get-ADGroup, New-ADUser, Set-ADUser, Add-ADGroupMember. File and folder operations: Get-ChildItem, Remove-Item, New-Item. Service management: Get-Service, Start-Service, Stop-Service. Event log queries: Get-EventLog, Get-WinEvent. Remote execution: Invoke-Command, Enter-PSSession. Script writing for administration tasks. Error handling and output formatting. Understanding PowerShell execution policies.
Practice Interview
Study Questions
Active Directory Architecture and Management
Active Directory structure: forests, domains, organizational units (OUs), users, groups, computers. Domain controllers and replication. LDAP and directory services concepts. Creating and managing user accounts, groups, and computer objects. User account properties, group membership, and delegation. Permissions on directory objects. Understanding global vs local groups, security vs distribution groups. Group nesting and scope (domain local, global, universal). Troubleshooting AD replication and connectivity issues.
Practice Interview
Study Questions
Infrastructure Design and Troubleshooting
What to Expect
90-minute technical interview with a senior infrastructure architect or lead. This round assesses your ability to think about infrastructure at a higher level beyond day-to-day operations. You'll receive real or realistic scenarios involving infrastructure design decisions, troubleshooting complex production issues, and making trade-offs. Expect questions like: 'Design a backup strategy for 100+ database servers', 'Walk through your approach to diagnosing complete service outage', 'How would you architect high availability for critical application', 'Troubleshoot performance degradation across 50-server cluster', 'Design strategy to migrate 200+ Linux servers to new data center'. This assesses infrastructure thinking, scoping complexity, and structured problem-solving.
Tips & Advice
For design questions, start by clarifying requirements: scale, availability requirements (RTO/RPO), budget constraints, complexity tolerance. Propose a solution, explain trade-offs (cost vs complexity, resilience vs simplicity), and discuss alternative approaches. For troubleshooting scenarios, use a systematic approach: scope the problem, establish timeline, gather metrics and logs, form hypotheses, test systematically, implement fix, verify resolution. Practice explaining infrastructure concepts to someone less technical—this is what you'll do mentoring junior staff and communicating with other teams. Study common infrastructure patterns: high availability, load balancing, database replication, caching layers, redundancy, failover strategies. Understand RAID concepts for storage. Be familiar with typical production outage causes and how to prevent them. For scenarios, ask clarifying questions about the environment, constraints, and success criteria. Discuss not just your solution but why you chose it over alternatives and what risks remain. Practice drawing infrastructure diagrams.
Focus Topics
Capacity Planning and Scalability
Assessing current infrastructure utilization and projecting growth. Determining when to scale resources. Vertical scaling (bigger servers) vs horizontal scaling (more servers) trade-offs. Predictive analytics for capacity planning. Cost modeling for infrastructure growth. Understanding scalability bottlenecks at different layers: database, application, network, storage.
Practice Interview
Study Questions
Infrastructure Documentation and Change Management
Maintaining accurate infrastructure documentation: architecture diagrams, runbooks, playbooks, configuration baselines. Change management processes: planning, testing, communicating, rolling back if needed. Version control for configuration files and scripts. Infrastructure as Code (IaC) concepts. Communication during outages and changes. Post-change validation.
Practice Interview
Study Questions
Monitoring, Alerting, and Observability
Designing comprehensive monitoring strategies: what to monitor, which metrics matter, alert thresholds, and avoiding alert fatigue. Understanding different types of alerts: threshold-based, anomaly detection, composite alerts. Monitoring tools and their use. Dashboards for visibility. Log aggregation and analysis for troubleshooting. Understanding latency, error rates, and business-relevant metrics alongside infrastructure metrics.
Practice Interview
Study Questions
Disaster Recovery and Business Continuity Planning
Recovery Time Objective (RTO) and Recovery Point Objective (RPO). Defining backup and recovery strategies based on business requirements. Full vs incremental vs differential backups. Backup location and off-site replication. Recovery procedures and testing. Automation of recovery processes. Understanding backup storage: local, network, cloud. Retention policies and compliance requirements. Database recovery strategies and point-in-time recovery.
Practice Interview
Study Questions
High Availability and Redundancy Architecture
Designing systems with high availability: active-active vs active-passive, load balancing, failover mechanisms, health checks. Understanding single points of failure and eliminating them. Database replication strategies (master-slave, master-master, multi-master). Application redundancy and stateless vs stateful services. DNS failover and traffic shaping. Geographic redundancy for disaster recovery. Practical tools: load balancers (software and hardware), clustering technologies, database replication.
Practice Interview
Study Questions
Production Incident Diagnosis and Root Cause Analysis
Systematic troubleshooting methodology for complex production issues: gather information quickly, establish baseline (what's normal), identify symptoms vs root cause, focus on most likely causes first, test hypotheses systematically, communicate status to stakeholders. Understanding common infrastructure failure modes: cascading failures, resource exhaustion, software bugs, network problems, hardware failures. Post-incident review and lessons learned. Preventing similar incidents through monitoring, alerting, and preventive measures.
Practice Interview
Study Questions
Backup, Disaster Recovery, and Business Continuity Specialization
What to Expect
60-minute focused technical interview on backup, disaster recovery (DR), and business continuity (BC) strategies. As a mid-level administrator in FAANG environment, you'll likely own backup and recovery operations for specific infrastructure. This round assesses your practical understanding of backup tools, recovery procedures, strategy design, and testing. Expect questions on RTO/RPO requirements, backup strategies for different workload types, recovery procedure verification, compliance requirements, and cost optimization of backup infrastructure.
Tips & Advice
Understand RTO (how quickly you need to recover) and RPO (how much data loss is acceptable)—these drive backup strategy. Study common backup architectures: full backups, incremental/differential backups, and hybrid approaches. Understand deduplication and compression for backup efficiency. Learn about backup tools (enterprise: NetBackup, Veeam, Commvault; open source: Bacula, rsync, Veeam Backup Free Edition). Know how to design backup strategies for different systems: databases require different approaches than file servers. Understand testing backup recovery—this is critical and often neglected. Be familiar with backup retention policies, archival tiers, and compliance requirements (GDPR, HIPAA, etc. retention rules). Study ransomware considerations: immutable backups, offline copies. Understand cloud backup and hybrid backup strategies. For databases, learn about transaction logs, full backups, and point-in-time recovery.
Focus Topics
Compliance and Retention Requirements
Understanding regulatory requirements: GDPR data retention, HIPAA backup requirements, SOX audit log requirements, industry-specific compliance. Impact on backup retention policies. Immutability requirements for certain backups (regulatory holds). Audit trails for backup operations. Encryption requirements for backups in transit and at rest.
Practice Interview
Study Questions
Database Backup and Recovery
Database backup strategies specific to database types (SQL Server, MySQL, PostgreSQL, Oracle, etc.). Full backups, transaction log backups, incremental backups. Point-in-time recovery. Backup consistency and transaction log management. Recovery time based on backup strategy. Using database-native tools (SQL Server backup, MySQL mysqldump, PostgreSQL pg_dump) and third-party backup tools. Backup verification queries to ensure completeness.
Practice Interview
Study Questions
Backup Tool Administration and Configuration
Practical experience with backup tools used in your environment (commercial or open source). Tool configuration: backup policies, schedules, retention rules. Monitoring backup job status and alert configuration for failures. Troubleshooting backup failures. Capacity planning for backup storage. Understanding tool-specific concepts: backup catalogs, media management, incremental forever vs traditional incrementals.
Practice Interview
Study Questions
Disaster Recovery Planning and Execution
Disaster Recovery Plan (DRP) creation and maintenance. Identifying Recovery Priority Objectives (RPO) for different systems. Tier 1 (critical business systems), Tier 2 (important systems), Tier 3 (nice to have). Failover procedures to DR site. Communication plans during disasters. Regular DR testing and drills. Documenting DR procedures in runbooks. Understanding RPO gap and acceptable data loss.
Practice Interview
Study Questions
Backup Strategy Design and Implementation
Designing backup strategies based on RTO/RPO requirements. Full vs incremental vs differential backups. Backup frequency and retention policies. Local vs offsite backups. Backup storage tiers (hot, cold, archive). Deduplication and compression. Cost-benefit analysis of backup strategies. Application-aware backups. Backup scheduling to minimize impact on production systems. Automation of backup processes.
Practice Interview
Study Questions
Recovery Procedures and Testing
Developing detailed recovery procedures for different failure scenarios. Step-by-step recovery runbooks. Testing recovery procedures regularly (at least quarterly). Tracking recovery testing metrics: MTTR (mean time to recovery), actual RTO achieved, RPO verification. Identifying gaps in recovery capabilities. Improving procedures based on test results. Documenting lessons learned from actual recovery situations.
Practice Interview
Study Questions
Behavioral and Leadership Principles
What to Expect
60-minute behavioral interview assessing cultural fit, leadership qualities, and how you embody company principles. For a mid-level role, this evaluates your ability to lead projects, mentor junior team members, collaborate cross-functionally, handle conflict, make decisions, learn continuously, and drive for results. Expect 5-7 behavioral questions using the STAR method (Situation, Task, Action, Result). Questions will likely probe: how you've led projects, handled ambiguity, solved problems collaboratively, dealt with failure, mentored others, advocated for your ideas, made tough trade-offs, and contributed to team decisions. This is assessed through concrete examples from your past.
Tips & Advice
Prepare 5-7 stories demonstrating key leadership and problem-solving behaviors, not generic company values. Use the STAR method: Situation (context, team size, what was the challenge), Task (your specific role and goal), Action (what you specifically did, decisions you made, how you led/collaborated), Result (quantified outcomes, what you learned). For mid-level, stories should show: project ownership (you led something end-to-end), mentorship (you helped junior team members grow), collaboration (you worked across teams to solve problems), decision-making (you made difficult trade-offs), and learning (you grew from failures). Avoid stories where you were just an individual contributor or follower. FAANG companies look for evidence of FAANG principles adapted to your level: Amazon's 'Ownership' (you owned outcomes, not just tasks), 'Learn and Be Curious' (you sought to understand root causes and improve), 'Earn Trust' (others relied on your judgment), 'Dive Deep' (you investigated issues thoroughly), 'Customer Focus' (your actions improved reliability/experience for users/teams), 'Think Big' (you contributed to strategy/planning). Prepare for common follow-up questions: 'What would you do differently?', 'What did you learn?', 'Who else was involved?'
Focus Topics
Decision-Making and Trade-offs
Stories showing you made difficult decisions with incomplete information, evaluated trade-offs, and justified your choice. Examples: technology selection, capacity planning decisions, prioritization of competing demands. Show how you gathered information, consulted others, made a decision, and communicated rationale.
Practice Interview
Study Questions
Collaboration and Cross-Functional Teamwork
Stories showing you working effectively with other teams: application development, database teams, network teams, security teams. Show how you aligned objectives, resolved conflicts, shared knowledge, and delivered joint solutions. Example: collaborating with dev team to troubleshoot production issue, working with database team to design backup strategy.
Practice Interview
Study Questions
Driving Results and Customer Impact
Stories where you improved infrastructure reliability, performance, or user experience. Show quantified results when possible: 'Improved backup recovery time from 4 hours to 30 minutes', 'Reduced manual operations by 70% through automation', 'Eliminated single point of failure impacting 500 users'. Show how you identified the problem, proposed solution, and delivered impact.
Practice Interview
Study Questions
Learning from Failure and Continuous Improvement
Stories showing you learned from mistakes or failures. What went wrong, how you responded, and what you learned. Examples: a deployment that went wrong (what did you learn to prevent it), a troubleshooting failure (how did you eventually solve it and improve your debugging skills), a tool selection that didn't work (why and what you'd do differently). Show growth mindset.
Practice Interview
Study Questions
Project Ownership and End-to-End Delivery
Demonstrating ability to own projects from conception to completion. Your story should show you taking responsibility for outcomes (not just tasks), defining requirements, making decisions, handling problems, and delivering results. Examples: implementing backup strategy for new infrastructure, leading migration project, deploying new monitoring system. Show how you coordinated across teams, managed timeline and scope, and achieved business objectives.
Practice Interview
Study Questions
Mentoring and Developing Junior Team Members
Show concrete examples of how you've helped junior or peer team members grow. This might include: teaching specific technical skills, providing feedback, delegating work with support, helping someone troubleshoot complex problems, involving them in design decisions. Show how your mentoring improved their capabilities and how you balanced autonomy with guidance.
Practice Interview
Study Questions
Hiring Manager Round
What to Expect
45-60 minute conversation with the hiring manager for the specific team. This is less formal assessment and more mutual evaluation. The manager assesses role fit, team dynamics compatibility, career aspirations, and ability to handle team-specific challenges. You learn more about the team, reporting structure, projects, challenges, and growth opportunities. Expect questions about your specific interests in the role, how you handle team dynamics, your career goals, and what kind of team environment you thrive in. This is your opportunity to ask substantive questions about the role and team.
Tips & Advice
Research the team and hiring manager if possible. Prepare 3-5 thoughtful questions about the role, team, and challenges. Questions might include: 'What are the biggest infrastructure challenges your team faces?', 'How does the team approach on-call and incident response?', 'What's your management style?', 'How do you support career growth?', 'What's the most interesting project the team is working on?' Come with genuine curiosity rather than a scripted approach. Be authentic about your career goals and what you're looking for in a role. Share why you're interested specifically in their team/company. Listen more than you talk—this is conversation, not presentation. If asked about team dynamics, give concrete examples: 'I work well with people who are collaborative and solution-focused', 'I appreciate direct feedback and try to give it constructively', 'I like working on teams with diverse perspectives'. Be honest about any concerns about the role or team—this is your chance to assess fit.
Focus Topics
Team Challenges and Problem-Solving Approach
Understanding team-specific challenges: recent outages, capacity constraints, technical debt, tool migrations, compliance requirements. How the team approaches problem-solving: data-driven, experimental, collaborative. Risk tolerance and change management philosophy.
Practice Interview
Study Questions
Career Development and Growth Opportunities
How the company supports career development. Opportunities for learning new skills and technologies. Mentorship availability. Path to senior roles or specialization. Support for certifications or training. Technical growth vs management track options.
Practice Interview
Study Questions
Team Dynamics and Working Style
How the team collaborates: meetings, communication channels, remote/office setup. Team size and structure. Interaction with other infrastructure teams and business teams. Support model: individual contributors, tech lead, manager. Conflict resolution approach. Team culture and values.
Practice Interview
Study Questions
Role and Team-Specific Responsibilities
Understanding the specific infrastructure domains, systems, and services this team owns. On-call rotation and incident response expectations. Day-to-day responsibilities vs strategic projects. Autonomy level and decision-making authority. How this role fits into broader infrastructure organization.
Practice Interview
Study Questions
Frequently Asked Systems Administrator Interview Questions
A small Linux web server serves only HTTPS to the public and takes SSH from a few admins. Design its host firewall policy, show the ruleset, and explain how you would test it and make it survive a reboot without cutting your own session.
Sample Answer
Direct answer
Use the kernel's own packet filter, nftables, on the server with a default-deny inbound policy: accept packets that belong to connections already allowed, accept loopback, accept a minimal set of ICMP, accept TCP 443 from anywhere, accept TCP 22 only from a named set of admin addresses, and drop everything else. Do not filter outbound for this host unless there is a reason (see the trade-off below). Test the file for syntax before applying it, apply it with a timed automatic rollback armed, confirm from a second, new session before cancelling the rollback, and make it permanent by putting it in the file the distribution's nftables service loads at boot.
If the host is already managed with ufw or firewalld, use that tool and express the same policy in it. Two firewall managers writing rules on one host is the failure to avoid.
Policy
| Traffic | Decision | Why |
|---|---|---|
| Inbound packets of an already-established or related connection | accept (first rule) | replies to the server's own outbound requests, and your existing SSH session, keep working while rules change |
Inbound packets in connection-tracking state invalid | drop | not part of any valid connection |
| Loopback interface | accept | local services talk to each other over it |
| ICMP and ICMPv6 error messages (destination-unreachable, time-exceeded, parameter-problem, packet-too-big) | accept | path MTU discovery (how hosts learn the largest packet that fits along the path) and error reporting fail without them. RFC 4890 section 4.3.1 says Destination Unreachable (all codes), Packet Too Big, Time Exceeded code 0 and Parameter Problem codes 1 and 2 must not be dropped, and section 4.3.2 says Time Exceeded code 1 and Parameter Problem code 0 normally should not be dropped; the ruleset accepts every code of all four types |
| ICMPv6 neighbour discovery (neighbour solicitation and advertisement, router advertisement) | accept | IPv6 hosts use them to find the other hosts on the link and the default gateway; blocking them breaks IPv6 on this host |
| Echo request (ping), IPv4 and IPv6 | accept, limited to 5 per second | lets monitoring and humans check reachability, with a cap on abuse |
| TCP 443 | accept from anywhere | the service this host exists to provide |
| TCP 22 | accept only from admin_v4 / admin_v6 sets | SSH exposed to the internet is the most attacked port; a handful of known addresses removes that exposure |
| Anything else inbound | drop (policy drop) | default deny: a service nobody planned for is not reachable by accident |
| Forwarding | drop | this is a server, not a router |
| Outbound | accept (see below) |
The firewall runs in the input hook of the server itself, so every packet addressed to the host passes it before any local program sees it. It complements, and does not replace, a cloud security group or a perimeter firewall: the host firewall is what still protects the server from other machines on the same network that have already been compromised.
The ruleset
#!/usr/sbin/nft -f
# Host firewall for a public HTTPS server with SSH from a few admin addresses.
# Replace only our own table, so rules from other tools (Docker, fail2ban) survive.
table inet filter
delete table inet filter
table inet filter {
set admin_v4 {
type ipv4_addr
elements = { 10.77.0.10, 10.77.0.11 }
}
set admin_v6 {
type ipv6_addr
elements = { fd00:77::10 }
}
chain input {
type filter hook input priority filter; policy drop;
ct state established,related accept
ct state invalid drop
iifname "lo" accept
# IPv6 does not work without neighbour discovery
icmpv6 type { nd-neighbor-solicit, nd-neighbor-advert, nd-router-advert, destination-unreachable, packet-too-big, time-exceeded, parameter-problem } accept
icmp type { destination-unreachable, time-exceeded, parameter-problem } accept
icmp type echo-request limit rate 5/second accept
icmpv6 type echo-request limit rate 5/second accept
tcp dport 443 accept
tcp dport 22 ip saddr @admin_v4 accept
tcp dport 22 ip6 saddr @admin_v6 accept
}
chain forward {
type filter hook forward priority filter; policy drop;
}
chain output {
type filter hook output priority filter; policy accept;
}
}
Reading the ruleset, line by line
| Line | What it means |
|---|---|
table inet filter | A table is a named container for chains and sets. The family inet makes one table cover IPv4 and IPv6 together; filter is only the table's name |
set admin_v4 { type ipv4_addr; elements = { ... } } | A named list of addresses. Rules refer to it as @admin_v4, so the allowed admins live in one place |
chain input { | A chain is an ordered list of rules. Packets are checked against the rules from top to bottom |
type filter hook input priority filter; policy drop; | type filter: this chain accepts or drops packets. hook input: attach it where packets addressed to this host arrive (forward is for packets passing through, output for packets the host sends). priority filter: the order among chains on the same hook; filter is the standard name for priority 0. policy drop: a packet that reaches the end without being accepted is dropped |
ct state established,related accept | ct is connection tracking, the kernel's memory of each flow. A packet that belongs to a connection already allowed (a reply, or a related error message) is accepted straight away, which is why this rule comes first |
ct state invalid drop | Packets that fit no known connection are dropped |
iifname "lo" accept | iifname is the input interface name. lo is the loopback interface, used by programs on this host talking to each other |
icmpv6 type { nd-neighbor-solicit, ... } accept | Accept the listed ICMPv6 message types: neighbour discovery (nd-...) plus the error types |
icmp type echo-request limit rate 5/second accept | Accept ping requests only while fewer than 5 per second arrive; the excess does not match this rule, falls through to the end and meets policy drop |
tcp dport 443 accept | dport is the destination port: anyone may connect to HTTPS |
tcp dport 22 ip saddr @admin_v4 accept | Port 22 and the source address (saddr) is in the set. The ip6 saddr line does the same for IPv6 |
chain forward / chain output | forward drops everything (not a router); output is policy accept, so outbound is unfiltered |
Each term in the policy table has a line that implements it:
| Policy row | Ruleset line |
|---|---|
| connection tracking (replies keep working) | ct state established,related accept and ct state invalid drop |
| path MTU discovery | packet-too-big in the icmpv6 line, destination-unreachable in the icmp line |
| neighbour solicitation and advertisement | nd-neighbor-solicit, nd-neighbor-advert, nd-router-advert in the icmpv6 line |
| SSH from admins only | the two tcp dport 22 ... saddr @admin_... lines |
Notes on the file:
- The first two lines of the table section (
table inet filterthendelete table inet filter) make the file re-loadable: the first creates the table if it does not exist (so the delete never fails), the second removes it, and the definition that follows rebuilds it. Only this table is replaced, so tables created by other tools (Docker's, fail2ban's) are left alone.flush rulesetwould wipe those too. I loaded the file onto an empty ruleset, then again onto the existing table, and it succeeded both times; a tableinet othercreated beforehand survived. - The family
inetcovers IPv4 and IPv6 in one table, so there is no forgotten IPv6 path with no rules. The two address sets are separate because the address types differ. The example addresses are a documentation-style private range (10.77.0.0/24) and a unique-local IPv6 address; replace them with your admin or bastion addresses. - Sets can be changed without reloading the file:
nft add element inet filter admin_v4 '{ 10.77.0.99 }'adds an address. To make it permanent, add it to the file as well, or the next reload forgets it. ct stateis connection tracking: the kernel remembers each flow, which is why one rule accepts all replies.
Test it before trusting it
-
Syntax check without applying:
nft -c -f /etc/nftables.confparses the file and reports errors without changing the running ruleset. -
Behaviour test from outside, using throwaway containers. This needs Docker on a Linux machine. Three containers share one Docker network: a server with the ruleset loaded and listeners on 22, 443 and 8080, one probe at an admin address and one at an address that is not in the set (the network, names and addresses are illustrative):
bash# setup, run from a directory holding web-fw.nft (the ruleset above), echo_server.py, server.sh and probe.sh docker network create --subnet 10.77.0.0/24 fw-test-net docker run -d --name fw-server --network fw-test-net --ip 10.77.0.2 --cap-add NET_ADMIN \ -v "$PWD":/w ubuntu:24.04 bash /w/server.sh docker run -d --name fw-admin --network fw-test-net --ip 10.77.0.10 -v "$PWD":/w ubuntu:24.04 sleep 300 docker run -d --name fw-other --network fw-test-net --ip 10.77.0.50 -v "$PWD":/w ubuntu:24.04 sleep 300 # wait for the server to finish installing and loading (about 30 seconds), then probe from each client docker exec fw-admin bash /w/probe.sh docker exec fw-other bash /w/probe.sh # cleanup docker rm -f fw-server fw-admin fw-other && docker network rm fw-test-net--cap-add NET_ADMINlets the server container change its own firewall; the clients do not need it. The echo server is a few lines of Python that listens on port 22 and sends back whatever it receives:python# echo_server.py import socket s = socket.socket() s.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1) s.bind(("0.0.0.0", 22)) s.listen() while True: c, _ = s.accept() while True: d = c.recv(100) if not d: break c.sendall(d)bash# server.sh (runs in the server container: Ubuntu 24.04) apt-get update -qq && apt-get install -y -qq nftables python3 > /dev/null python3 /w/echo_server.py & # echo server on TCP 22 for p in 443 8080; do python3 -m http.server "$p" --bind 0.0.0.0 > /dev/null 2>&1 & done nft -c -f /w/web-fw.nft && nft -f /w/web-fw.nft wait # keep the container (and the listeners) alive # probe.sh (runs in each client container; the server is 10.77.0.2) for p in 22 443 8080; do timeout 2 bash -c "exec 3<>/dev/tcp/10.77.0.2/$p" 2>/dev/null \ && echo "port $p open" || echo "port $p blocked/timeout" doneReading the probe:
/dev/tcp/HOST/PORTis not a real file but a bash feature, andexec 3<>/dev/tcp/10.77.0.2/22makes bash try to open a TCP connection to that address and port as file descriptor 3. If the connection is accepted the command succeeds, so&&prints "open". A dropped packet gets no reply, so the attempt hangs untiltimeout 2kills it after 2 seconds and||prints "blocked/timeout". A closed port that is not firewalled would answer at once with a refusal and print the same blocked message. Without the finalwaitthe script ends, the container exits and every probe prints blocked/timeout. With it, the probes printed exactly the table below.Probe from port 22 port 443 port 8080 admin address 10.77.0.10 open open blocked/timeout other address 10.77.0.50 blocked/timeout open blocked/timeout This confirms all three decisions: the public service is reachable by everyone, SSH only by the admin set, and an unplanned port by nobody.
-
Check that reloading does not drop a live session. A client at an admin address opens a TCP session to port 22, the server's ruleset is reloaded with
nft -fwhile the session is open, and the client sends again:python# session_client.py <server-ip> import socket, sys, time s = socket.create_connection((sys.argv[1], 22), timeout=5) s.sendall(b"before"); print("echo 1:", s.recv(100)) time.sleep(6) # the server reloads its ruleset during this pause s.sendall(b"after-reload"); print("echo 2:", s.recv(100))Output:
echo 1: b'before'thenecho 2: b'after-reload'. The established-connection rule is what carries the session across the reload.
Applying it on a real host without cutting yourself off
Two ways lock-outs happen: a typo in the admin address set, and rules that load fine but block the path you are using. Guard against both:
# 1. save our own table so it can be put back (or removed again if it did not exist)
{ echo 'table inet filter'; echo 'delete table inet filter'; nft list table inet filter 2>/dev/null; } > /root/fw-before.nft
# 2. arm a rollback in 2 minutes (systemd transient timer)
systemd-run --unit=fw-revert --on-active=2min /usr/sbin/nft -f /root/fw-before.nft
# 3. validate, then apply
nft -c -f /etc/nftables.conf && nft -f /etc/nftables.conf
# 4. from a SECOND terminal, open a new SSH connection from an admin address
# and request the HTTPS page from outside; only then cancel the rollback:
systemctl stop fw-revert.timer
Reading step 1: the curly braces form a brace group, which runs the three commands one after the other and lets one > send their combined output to a single file. The file therefore starts with table inet filter (create the table if it is missing, so the next line cannot fail) and delete table inet filter (remove it), followed by nft list table inet filter (the current definition of that table in loadable form, or nothing if there is no such table). Loading the file later first removes our table and then re-creates the old one if there was one. Only our own table is touched, which matches how the ruleset file itself is loaded. In step 2, systemd-run --on-active=2min schedules that load to run once, 2 minutes from now.
If the new rules strand you, you do nothing and the timer restores the old behaviour in 2 minutes. The saved file begins with the create-then-delete pair because loading a file does not remove rules that are already present; I checked that loading an empty file left the new rules in place, that the saved file removed the table again when none existed beforehand, and that, when one did exist, adding a rule and then loading the saved file brought back exactly the saved rules. I did not save the whole ruleset with flush ruleset because that is not safe on a host that also runs Docker: nft list ruleset prints the iptables-nft compatibility tables too, and in a test container nft -f refused to load them back with unsupported xtables compat expression. A file loads as one transaction, so that error would have meant the revert changed nothing and the lock-out stayed in place, and a successful flush would also have wiped the other tools' tables that the ruleset file takes care to leave alone. systemd-run --on-active= creates a transient timer, and with --unit= the timer is named fw-revert.timer; the revert command itself was not run here, because the test containers have no systemd. Keep out-of-band access (a cloud serial console or hypervisor console) ready, since it is the only route back if even the timer fails.
Surviving a reboot
On Debian and Ubuntu the nftables package ships nftables.service, whose ExecStart is /usr/sbin/nft -f /etc/nftables.conf, so the file above goes in /etc/nftables.conf and systemctl enable nftables makes it load at boot. On the RHEL family the service loads /etc/sysconfig/nftables.conf instead (in the AlmaLinux 9 package, whose ExecStart is /sbin/nft -f /etc/sysconfig/nftables.conf, that file ships with the sample include commented out), so use include there or put the rules in that file. To test persistence without waiting for a reboot, run nft flush ruleset from a console session and then systemctl restart nftables and list the ruleset; then do one real reboot in a maintenance window, with console access open, before calling it done.
Trade-offs and what would change the design
- Outbound
accept: a web server that can reach anywhere makes data theft and reverse shells easier after a compromise. The alternative is an outbound allow-list (DNS, time sync, package repositories, the specific APIs the application calls). I would take that for a server holding sensitive data, and accept the maintenance cost (every new dependency is a ticket). For a small server whose compromise has limited blast radius, accept-all outbound plus egress monitoring is a defensible choice. - SSH allow-list versus a bastion or VPN: if admins have changing addresses, an allow-list becomes constant churn and lock-outs. Allow SSH only from the bastion's address instead, and control who can use the bastion.
- Rate-limiting echo and new connections is a convenience, not a security boundary. Real denial-of-service protection happens upstream.
- Containers on the same host (Docker, Kubernetes) add their own rules and forwarding behaviour. The scoped
delete tableapproach avoids deleting them, but you must test the combination.
Pitfalls
- Putting
ct state established,related acceptafter the SSH rule, or omitting it, and cutting off the session when the policy loads. - Forgetting IPv6: an IPv4-only ruleset leaves the host wide open over IPv6 if it has an address, which is why the table is
inet. - Setting
policy dropon the input chain before confirming the ICMPv6 neighbour-discovery rules, so the host loses its default gateway and drops off the network. - Testing only from inside the host, which shows nothing about what the network can reach.
Deep specialization in one area versus staying a broad generalist: which would you choose for your own career from here, and what are you consciously trading away?
Sample Answer
Direct answer
Neither path is inherently better. The honest answer names what you're optimizing for right now, depth of leverage and marketability in a narrow area, versus flexibility and broader career options, and states plainly what you're giving up by choosing one, rather than pretending you can maximize both at once.
Structured elaboration
Define the trade-off in your own terms. Deep specialization trades breadth of future options for concentrated leverage and recognition in one area. Staying a broad generalist trades peak depth in any one area for flexibility, resilience to shifts in what your organization needs, and often a more natural path into roles that require breadth.
| Dimension | Deep specialist | Broad generalist |
|---|---|---|
| Leverage | Concentrated impact within one domain | Cross-cutting impact connecting systems or teams |
| Marketability | Strong where that specific depth is valued, narrower market | Broader market, easier lateral moves |
| Risk | Exposure if the narrow area loses relevance | Risk of shallow expertise without a differentiated edge |
| Typical path | Domain authority, principal-track recognition | Leadership, architect, or cross-functional roles |
Name what you're consciously trading away, specifically. If you specialize, you accept slower or harder pivots later and reliance on organizations that value that specific depth. If you generalize, you accept giving up the strongest, most differentiated reputation in any single area, and possibly slower recognition in fast, depth-rewarding tracks.
Ground the choice in something real. Your current stage, early career often benefits from some depth to build a track record, later career often benefits from breadth for leadership options, what your organization or market currently rewards, and where your genuine interest sustains itself over time.
Apply a useful test. Describe a specific moment where you actually had to choose between a deep technical option and a broader, stakeholder-facing one, and what you picked. A real decision under real constraint tells an interviewer far more than a stated preference in the abstract.
Worked example
"At one point I had two real options in front of me at the same time, a deep technical project that would make me the clear expert in a narrow area few others touched, or a stakeholder-facing initiative that would put me in front of more of the organization with less technical depth involved. I chose the stakeholder-facing option, consciously, because at that stage I already had reasonable depth in my area and what I was missing was visibility and cross-functional experience, which the deep project wouldn't have given me regardless of how well I executed it. I was explicit with myself that I was trading a chance to become the clear go-to expert in that narrow area for broader relationships and exposure, and that someone else would likely become that expert instead. Looking back, the choice matched what that stage of my career actually needed, which is the test I'd apply again, not which option sounds more impressive, but which trade-off fits where I am now."
Trade-offs & pitfalls
- Treating this as a values statement, I love learning new things, without naming the actual cost of the choice reads as avoiding the harder half of the question.
- Claiming you can do both fully at once. Some blending is real, build depth then broaden, or vice versa, in phases, but pretending there's no trade-off undercuts your credibility.
- Answering based on what sounds better in an interview rather than what you'd actually choose usually shows in the lack of a concrete supporting example.
- A generalist claim with no depth anywhere reads as avoiding commitment, just as a specialist claim with no awareness of the narrowing risk reads as naive about the market.
How can client-side and server-side Git hooks or pre-commit tools be used to improve the safety of infrastructure changes? Give examples of checks you'd run locally and checks you'd run in CI for an infrastructure repository, listing specific tools where appropriate.
Sample Answer
Direct answer
Git hooks (client-side, running on a developer's own machine before a commit or push completes) and CI checks (server-side, running centrally on every push regardless of the developer's local setup) are COMPLEMENTARY, not substitutes: client-side hooks give the fastest possible feedback (catch a mistake before it is even committed, saving a round trip), while CI is the only place a check can be TRUSTED as actually having run, since a client-side hook can always be skipped (--no-verify) or simply not installed on a given machine. A safe design runs the SAME class of checks in both places, fast local hooks for immediate feedback, the identical (or stricter) checks re-enforced in CI as the actual gate that matters.
Structured elaboration
Client-side, pre-commit. Runs locally before a commit is created. Good fit for FAST, deterministic checks with no external dependencies: terraform fmt -check (formatting), a YAML/JSON syntax lint, a check for accidentally-staged files matching a secret-pattern regex (a cheap, local first line of defense, not a substitute for a real secret scanner run in CI), and terraform validate if the local environment has provider plugins already available. Tools: the pre-commit framework (language-agnostic hook management, widely used specifically for infra repos), or Terraform-specific wrappers like tflint run as a pre-commit hook.
Client-side, pre-push. Runs before a push reaches the remote, a natural place for slightly heavier checks that are still local-only (a broader tflint rule set, a local policy-as-code dry run against cached provider schemas) that would be too slow to run on every single commit but are still worth catching before code leaves the developer's machine.
Server-side, in CI. Runs on every push to the remote, REGARDLESS of whether the developer had hooks installed, skipped them, or is pushing from an environment (a different machine, a bot) that never had local hooks configured at all. This is where the checks that actually GATE merge belong: terraform plan (requires real provider credentials, cannot run meaningfully as a lightweight local hook for most teams), a real secret scanner (gitleaks or similar, scanning full history and diff context, not just a local regex), policy-as-code evaluation (OPA or Sentinel against the plan output), and the full test suite (terratest/kitchen-terraform).
Server-side, as a Git-platform-native server hook (less common, but worth naming). Some Git hosting platforms support genuine SERVER-SIDE hooks (distinct from CI, running as part of the push itself, able to REJECT a push outright before it is even accepted) for very fast, universal checks, a commit-message format check, or a hard block on pushing directly to a protected branch; where available, this is stronger than CI for checks that should be impossible to bypass even accidentally, since CI only runs AFTER a push has already succeeded.
Worked example
A concrete layered setup for an infrastructure repository:
| Layer | Checks | Tooling |
|---|---|---|
| Pre-commit (local) | terraform fmt -check, YAML/JSON lint, cheap secret-pattern regex | pre-commit framework |
| Pre-push (local) | tflint (broader ruleset), terraform validate | pre-commit framework, push stage hooks |
| Server-side push hook (platform-native, if available) | Commit message format, block direct pushes to main | Git hosting platform's native server hooks |
| CI (server-side, authoritative) | terraform plan, full secret scan (gitleaks), policy-as-code (OPA), terratest suite | CI pipeline, required status checks |
Trade-offs and pitfalls
- Common mistake: treating a passing local pre-commit hook as sufficient confidence to merge, without the SAME checks also running, and gating, in CI. A hook is trivially bypassable (
git commit --no-verify, or simply never having installed it), so any check whose absence would be a real problem MUST also run server-side as a required status check; local hooks are a speed optimization for the common case, never the actual safety guarantee. - Putting genuinely slow or credential-requiring checks (a full
terraform planagainst real infrastructure) into a pre-commit hook creates enough friction that developers start reaching for--no-verifyhabitually, once that becomes normal behavior for ANY check, it quietly becomes normal for ALL checks, eroding the whole local-hook layer's value; keep local hooks fast and dependency-free specifically to protect this. - A cheap local secret-pattern regex catches obvious accidental staging of an API key but is not a substitute for a real secret scanner in CI, which checks full history and diff context and is maintained against known secret-format patterns far more broadly than a team's own local regex is likely to cover.
- Server-side, platform-native hooks (not CI, but true push-time hooks) are underused specifically because fewer teams know their Git hosting platform supports them, for checks that should be genuinely impossible to bypass (not just gated on merge, but rejected at push time), this is a stronger tool than CI and worth confirming platform support for before assuming CI is the only server-side option.
Schedule a nightly ETL job to run at 02:30 with cron. Provide the crontab line, explain its fields, and describe why a script that works in your terminal can fail under cron and how you would make failures visible and prevent overlapping runs.
Sample Answer
Direct answer. The crontab line is 30 2 * * * /opt/etl/run_etl.sh >> /var/log/etl/nightly.log 2>&1. The five fields are minute, hour, day of month, month and day of week, so 30 2 * * * means minute 30 of hour 2 (02:30, 24-hour clock) every day. A script that works in your terminal can fail under cron (the daemon that runs scheduled jobs) because cron starts it with a different environment: a minimal PATH, no profile files, a different working directory and /bin/sh as the shell. Make failures visible by logging with timestamps, exiting non-zero on failure, and alerting when the job did not succeed. Prevent overlapping runs with a lock, flock -n (flock takes a lock on an open file, -n means fail at once instead of waiting), so a run that takes longer than expected is not started twice.
The fields
| Field | Allowed | In 30 2 * * * |
|---|---|---|
| Minute | 0 to 59 | 30 |
| Hour | 0 to 23 | 2 |
| Day of month | 1 to 31 | * (every day) |
| Month | 1 to 12 | * |
| Day of week | 0 to 7 (0 and 7 are Sunday) | * |
After the five time fields comes the command, run by /bin/sh unless the crontab sets SHELL=. ETL (extract, transform, load) is the pipeline that copies data from sources into a warehouse. Cron times use the system clock's time zone: run servers in UTC, because on days when clocks change the 02:30 slot may not exist or may occur twice, and cron implementations differ in what they do about it.
Why it works in a terminal but not under cron
In a clean Ubuntu 24.04 container with cron installed, a probe job scheduled every minute logged the environment cron really provides:
SHELL=/bin/sh PATH=/usr/bin:/bin HOME=/root PWD=/root
- PATH has no
/usr/local/binor/sbin, so tools installed there (and anything added by your profile,nvm, a virtualenv or~/.local/bin) are "command not found" (status 127). - No profile.
.bashrcand.profileare not read, so exported credentials and variables are absent. - Working directory is the home directory, so relative paths break. Use absolute paths or
cdfirst. - Shell is
/bin/sh, so bashisms ([[ ]], arrays,pipefail) fail unless the script has a#!/usr/bin/env bashline (when the script is run directly rather than as an argument tosh) or the crontab setsSHELL=/bin/bash. - Percent signs are special in a crontab. An unescaped
%ends the command and the rest becomes standard input. Writingecho "stamp $(date +%F)"in a crontab produced no log file at all in that test, because the shell received a truncated, invalid line.$(date +\%F)worked and loggedstamp 2026-10-06. Put the command in a script and call that. - No terminal and no interactive prompts. Output goes to mail (if configured) or is lost, so redirect it.
A wrapper that fixes the environment, avoids overlap and reports failure
#!/usr/bin/env bash
# /opt/etl/run_etl.sh : wrapper cron calls. Makes the environment explicit.
set -euo pipefail
export PATH=/usr/local/bin:/usr/bin:/bin # do not rely on cron's PATH
cd /opt/etl # do not rely on cron's cwd
log() { printf '%s nightly-etl: %s\n' "$(date '+%F %T')" "$*"; }
# One run at a time. Descriptor 9 holds the lock until this process exits.
exec 9> /var/lock/nightly-etl.lock
if ! flock -n 9; then
log "previous run still holds the lock, skipping this one"
exit 75 # EX_TEMPFAIL: not an ETL failure
fi
on_error() { log "FAILED at line $1 (exit $2)"; exit "$2"; }
trap 'on_error "$LINENO" "$?"' ERR
log "start"
etl-tool --extract # a command that lives in /usr/local/bin
sleep "${ETL_SECONDS:-1}" # stands in for the real transform and load
log "done"
touch /var/lib/etl-last-success
Reading the wrapper line by line:
-
exec 9> /var/lock/nightly-etl.lockopens the lock file for writing and keeps it open as file descriptor 9. A file descriptor is just a small number a process uses to refer to a file it has open (0, 1 and 2 are standard input, output and error already, so 9 is a free number).execwith only a redirection changes the current shell instead of starting a program, so the file stays open until the script exits. Opening the file does not lock it; the next line does. -
flock -n 9asks the kernel for an exclusive lock on whatever descriptor 9 points to. If another process already holds that lock the command fails immediately, and theif !branch logs and exits. Measured in a container (readlink /proc/$$/fd/9shows what descriptor 9 points to):/tmp/demo.lock first: got the lock second: busy (flock exit 1) third: got the lockThe first shell opened the file on descriptor 9 and locked it, a second process that tried to lock the same file was refused, and after descriptor 9 was closed (
exec 9>&-) a third process got the lock. When a script exits, the kernel closes its descriptors, and the lock is released when the last open copy of descriptor 9 is closed. Children the wrapper starts inherit descriptor 9, so a child that outlives a killed wrapper keeps the lock held (executed: afterkill -9of a shell holding the lock, a still-runningsleepchild madeflock -non the same file fail, while a child started with9>&-did not). That is usually what you want, because the job is still running; start any helper that must not hold the lock with9>&-. -
exit 75is the status for a skipped run. 75 isEX_TEMPFAIL("temporary failure, try later") from the BSDsysexits.hlist of conventional exit codes, so a monitor can tell a skip from a real failure (any other non-zero status). -
trap 'on_error "$LINENO" "$?"' ERRregisters a handler for commands that fail whileset -eis on. The single quotes matter:$LINENO(the line of the failing command) and$?(that command's exit status) are only filled in when the trap fires, not when it is registered. Measured on this toy script (saved as a file, with the shebang as line 1), which registers the same trap and then runsfalse, a command that always fails:bash#!/usr/bin/env bash set -euo pipefail on_error() { echo "FAILED at line $1 (exit $2)"; exit "$2"; } trap 'on_error "$LINENO" "$?"' ERR echo start false echo not reachedstart FAILED at line 6 (exit 1) wrapper exit status: 1(the line number counts from the top of that toy file, so it differs from the 20 reported for the real wrapper below.)
-
${ETL_SECONDS:-1}means "the value ofETL_SECONDSif it is set and not empty, otherwise1":sleepwaits 1 second normally, and the lock test below sets it to 75.
The crontab entry calls the wrapper and appends both streams to a log:
30 2 * * * /opt/etl/run_etl.sh >> /var/log/etl/nightly.log 2>&1
To check the lock, the wrapper was run from real cron every minute with the job taking 75 seconds (ETL_SECONDS=75):
2026-10-06 06:06:01 nightly-etl: start
extracted
2026-10-06 06:07:01 nightly-etl: previous run still holds the lock, skipping this one
2026-10-06 06:07:16 nightly-etl: done
The 06:07 run found the lock held and skipped itself instead of starting a second copy. The kernel releases the lock once no process holds the descriptor open any more, even if the holder was killed, so there is no stale lock file to clean up. With the extractor replaced by a stand-in that prints an error and exits 3, the wrapper logged FAILED at line 20 (exit 3), exited with status 3 and did not create the success marker.
Making failures visible
- Log with timestamps and keep stderr in the log (
2>&1), as above. - Exit non-zero on failure so whatever supervises the job can see it. The wrapper uses
set -euo pipefailand an ERR trap. - Alert on absence as well as failure. If cron itself is down or the host is off, no error is ever logged. Record a success marker (the wrapper touches
/var/lib/etl-last-success) and run a separate check that alerts when it is older than your threshold (for example 26 hours for a daily job), or ping an external heartbeat service at the end of a successful run. - Mail:
MAILTO=ops@example.comat the top of the crontab sends any output the job produces to that address, which needs a working mail transfer agent (MTA, the program on the host that actually delivers email, such as Postfix) on the host.
Trade-offs and pitfalls
- Skipping versus queuing.
flock -nskips the overlapping run. If the data must be processed, run the next job withflockwithout-nso it waits, but then a stuck run blocks every night after it. A timeout (flock -w 600) is a middle ground. - A skipped run is not a failure, so the wrapper exits 75 (the BSD
EX_TEMPFAILconvention) to keep it distinguishable from a real error. - Idempotency (running twice gives the same result as running once). Cron can run the job twice (manual rerun, host time change), so the ETL should be safe to repeat, for example by loading by date partition.
- systemd timers are an alternative: a service unit is not started again while it is still running, output goes to the journal, and
OnFailure=can trigger an alert. Cron is fine for one nightly job when it has the wrapper above. - Cron does not retry. A failure at 02:30 is not retried until the next day unless you build it in.
Explain what duplicate ACKs mean and how they trigger TCP's fast retransmit, without waiting for the retransmission timer to fire. How would you tell, from the pattern of duplicate ACKs and retransmissions, whether the cause is genuine packet loss versus packet reordering along the path, and why does that distinction change which mitigation is appropriate?
Sample Answer
Direct answer
A duplicate ACK is the receiver re-acknowledging the same byte offset it already acknowledged, which happens when a later, out-of-order segment arrives before the one actually missing; three duplicate ACKs in a row is treated as strong enough evidence of real loss to trigger an immediate fast retransmit, without waiting for the slower retransmission timer.
Structured elaboration
Normally, each ACK acknowledges progressively more data as segments arrive in order. If segment N is lost but segment N+1 arrives (out of the expected order relative to what the receiver is still waiting for), the receiver can't advance its cumulative ACK past the end of segment N-1, so it re-sends an ACK for the SAME byte offset it already acknowledged, that's the duplicate ACK. One or two duplicate ACKs are common and unremarkable (ordinary, brief reordering happens on real networks); but THREE duplicate ACKs in a row is treated as a strong enough signal that data is genuinely missing (not just briefly reordered) to justify retransmitting immediately, well before the (much slower) retransmission timeout would otherwise fire.
Distinguishing genuine loss from simple reordering, in practice: sustained duplicate ACKs (three or more, continuing as MORE out-of-order segments keep arriving) point to real loss, since a purely reordering event typically self-resolves within one or two duplicate ACKs once the delayed segment catches up. A pattern of isolated single or double duplicate ACKs that stop on their own, with no retransmission ultimately needed, points to reordering rather than loss; TCP's own reordering-tolerance thresholds (adjacent to, but distinct from, PAWS: Protect Against Wrapped Sequence numbers) exist specifically to avoid triggering a fast retransmit on every minor reordering event.
Worked example
A sender transmits segments carrying bytes 1000-1500, 1500-2000, 2000-2500, and 2500-3000. If the segment carrying 1500-2000 is lost but the other three arrive, the receiver sends: ACK 1500 (for the first segment, normal), then ACK 1500 again upon receiving 2000-2500 (a duplicate, since 1500-2000 is still missing), then ACK 1500 again upon receiving 2500-3000 (a second duplicate). On the THIRD duplicate ACK for 1500, the sender's fast retransmit fires and it resends the 1500-2000 segment immediately, rather than waiting for its retransmission timer (which, per the RTO calculation, could be tens to hundreds of milliseconds longer) to expire.
Trade-offs & pitfalls
The right mitigation depends entirely on which cause is confirmed: if it's genuine loss, the useful levers are addressing the actual loss source (a congested link, a flaky physical connection) or, at the transport layer, ensuring SACK (Selective Acknowledgment) is enabled so only the truly missing segment gets resent. If it's reordering (for instance, from a load balancer or ECMP path hashing packets across multiple physical paths with slightly different latencies), the fix is architectural (favor flow-based hashing that keeps a single connection's packets on one path) rather than anything TCP-level, since TCP is already tolerating ordinary reordering correctly; treating a reordering pattern as loss and "fixing" the transport layer for it addresses the wrong layer.
A Windows Server feels slow. Which performance counters would you look at to tell whether the bottleneck is CPU, memory, disk or network, what values would worry you, and how do you know what normal looks like for that server?
Sample Answer
Direct answer
Look at one counter per resource, and always read them together: CPU (\Processor(_Total)\% Processor Time), memory (\Memory\Available MBytes), disk (\PhysicalDisk(_Total)\Avg. Disk sec/Read and Avg. Disk sec/Write, the average time per I/O, where one I/O is a single read or write request to the disk, plus Current Disk Queue Length, the number of requests waiting for the disk) and network (\Network Interface(*)\Bytes Total/sec against the link speed). Worrying values depend on the server's role, so the real test is deviation from that server's own baseline: a recorded picture of normal load, captured with a Data Collector Set (a saved list of counters, a sample interval and an output file that Performance Monitor can start and stop) or logman over at least a normal business cycle. Task Manager and Resource Monitor are the quick first look to see which resource and which process; counters and a baseline tell you whether it is abnormal.
One counter per resource, and what worries me
| Resource | Counter | What I look for |
|---|---|---|
| CPU | \Processor(_Total)\% Processor Time | Sustained high use, not spikes. Microsoft's Windows performance troubleshooting guide rates user-mode CPU (\Processor Information(*)\% User Time) above 80% as critical (50 to 80% is a warning) and % Idle Time under 10% as critical, and says to start investigating when it stays high for a minute or more. % Processor Time is user plus kernel time, so use those figures as a guide rather than a limit. |
| Memory | \Memory\Available MBytes | A low and falling value. The same guide rates it healthy at over 10% free (or at least 4 GB) and critical below 1% or below 500 MB; a database or file server will have a different normal. |
| Disk | \PhysicalDisk(_Total)\Avg. Disk sec/Read and /Write | Latency per I/O in seconds (0.015 is 15 ms). The Windows performance guide rates under 15 ms healthy, over 25 ms a warning and over 50 ms critical, and says to investigate latency that lasts a minute or longer; Microsoft's Hyper-V tuning guidance likewise acts when host disk latency is consistently over 50 ms. Check Current Disk Queue Length to see whether I/O is piling up. There is no single published queue number that applies to every disk (a RAID set or SAN volume serves many requests at once), so judge the queue against this server's baseline and together with latency: a queue that grows while latency rises means requests are waiting. |
| Network | \Network Interface(*)\Bytes Total/sec | Sustained use near the link capacity. The Windows performance guide rates under 50% of the NIC's speed healthy and over 80% critical; the Hyper-V guidance acts at 90% of the physical NIC's capacity. |
These figures are published starting points for particular workloads, not universal limits: a batch server that runs hot at 95% CPU overnight is healthy; a file server at 60% CPU may not be. If a counter path is not found on your build, list the exact names with (Get-Counter -ListSet 'Network Interface').Counter.
Telling the bottleneck apart
Read the counters as a pattern, not one by one:
- CPU-bound:
% Processor Timehigh, disk latency normal, memory normal. Then find the process in Task Manager or Resource Monitor. - Memory-bound:
Available MByteslow and falling while disk latency and disk activity rise, because Windows starts paging to disk (when RAM runs short it writes the least-used memory pages to the pagefile on disk and reads them back later, which creates disk traffic that has nothing to do with your files). Disk looks guilty but memory is the cause. - Disk-bound: latency high with a growing queue while CPU is low (processes are waiting on I/O).
- Network-bound: bytes per second near the link capacity, with CPU and disk idle. Slow responses with low use everywhere point to something outside the box (a database, DNS or a remote dependency).
How you know what normal looks like
Record a baseline before there is a problem. A 15 second sample interval is enough for trends; a one second interval is for live diagnosis.
# Quick live look: 5 samples, 2 seconds apart, four resources at once
Get-Counter -Counter @(
'\Processor(_Total)\% Processor Time',
'\Memory\Available MBytes',
'\PhysicalDisk(_Total)\Avg. Disk sec/Read',
'\PhysicalDisk(_Total)\Avg. Disk sec/Write',
'\PhysicalDisk(_Total)\Current Disk Queue Length',
'\Network Interface(*)\Bytes Total/sec'
) -SampleInterval 2 -MaxSamples 5
# Long-running baseline: a circular binary log, 15 s samples, opened later in Performance Monitor.
# -c counters to collect, -si 15 sample every 15 seconds, -f bincirc binary circular log format
# (once the file reaches the -max size it overwrites its oldest data), -max 512 maximum file size in MB,
# -o output path.
logman create counter ServerBaseline -c "\Processor(_Total)\% Processor Time" "\Memory\Available MBytes" "\PhysicalDisk(_Total)\Avg. Disk sec/Read" "\PhysicalDisk(_Total)\Avg. Disk sec/Write" "\Network Interface(*)\Bytes Total/sec" -si 15 -f bincirc -max 512 -o C:\PerfLogs\ServerBaseline
logman start ServerBaseline
Reading the live look: Get-Counter prints one block per sample, a Timestamp and then each counter path followed by its value. Illustrative abridged output (the numbers are examples, and the values for this server will differ):
Timestamp CounterSamples
--------- --------------
10/5/2026 09:00:02 \\fs01\processor(_total)\% processor time :
19.8
\\fs01\memory\available mbytes :
5120
\\fs01\physicaldisk(_total)\avg. disk sec/read :
0.045
\\fs01\physicaldisk(_total)\current disk queue length :
14
Here CPU is 19.8 percent, 5,120 MB of memory is available, reads take 0.045 s (45 ms) each and 14 requests are waiting for the disk. Compare each value to the baseline: only the disk values stand out. The counter names are lower-cased in the output; that is normal.
Capture across a full business cycle (at least a week if there is a weekly pattern, and include month-end if the server has one), then record the median and the busiest hour for each counter. Alert on deviation from those numbers.
Worked example
A file server reports "slow" every Monday at 09:00. The baseline shows Avg. Disk sec/Read normally around 0.004 (4 ms). During the complaint the quick sample shows 0.045 (45 ms), CPU at 20% and Available MBytes steady. Latency is more than ten times its baseline (45 / 4 = 11.25) while CPU and memory are normal, so the bottleneck is storage. A 1 Gbps link carries 1,000,000,000 / 8 = 125,000,000 bytes per second, so 90% of capacity is 112,500,000 bytes per second; the NIC counter reads far below that, which rules out the network. Resource Monitor's Disk tab then shows which process and file are generating the reads (a Monday scheduled backup scan).
Trade-offs and pitfalls
- Never judge one counter in isolation. High disk latency with low memory is a memory problem; high CPU with a growing disk queue may be CPU waiting on slow I/O.
- Averages hide spikes. Keep the sample interval small enough for the symptom (a 15 second sample will miss a two-second stall).
- A baseline taken during the incident is not a baseline. Capture it while the system is healthy and keep it as a reference.
- Virtual machines: the guest's counters can look fine while the host is overcommitted, so check host-side counters as well.
You must present the postmortem for a significant outage to non-technical executives, and potentially to customers or the public. How does the structure and level of detail change from the internal engineering postmortem? Describe what you include and omit, how you present root cause and remediation without minimizing real impact, and how you handle information that is sensitive or under legal review.
Sample Answer
Direct answer
An executive or public postmortem communication keeps the same underlying facts as the internal engineering document but changes structure and depth: lead with impact and resolution status in plain language, compress the technical root cause into one or two sentences a non-specialist can follow, and route anything sensitive, legally uncertain, or still under investigation through legal or compliance review before it goes out, rather than including it by default.
Structured elaboration
- Lead with what the audience actually needs. Executives and customers care first about impact (who was affected, how badly, for how long) and current status (is it fixed, is it safe now), not the internal technical mechanism. Put that first, not buried after a long technical narrative.
- Compress, don't omit, the root cause. A one or two sentence plain-language root cause ("a configuration change removed a safeguard that normally limits how much traffic a single request can trigger") is usually enough; the internal document's full technical detail isn't needed here and can overwhelm or confuse rather than reassure.
- Say what's being done, concretely. Vague reassurance ("we take this seriously and are reviewing our processes") reads as evasive. Specific, verifiable commitments ("we are adding an automated safeguard, expected within two weeks") build more trust even when the news is bad.
- Route sensitive content through review before drafting is even final. Anything touching legal exposure, an ongoing investigation, regulatory disclosure requirements, or third-party or customer data (for example a possible PII exposure) needs legal or compliance sign-off on both content and timing, since public/customer communication commitments here can create legal exposure of their own if stated imprecisely.
- Don't minimize real impact to make the story feel better. Understating severity or hedging around clear facts, once discovered (and it usually is), costs far more trust than a direct, honest account would have.
This same discipline extends past software outages: a public account of a failed research study that led to a wrong decision, or a partnership failure with a strategic account, follows the identical shape (impact first, plain-language cause, concrete next steps), adapted in vocabulary but not in structure.
Worked example
An internal postmortem for a data-exposure incident runs several pages with full technical detail about the specific misconfigured storage permission, exact timestamps, and internal system names. The customer-facing version: a short notice stating what data was potentially exposed (in plain terms, not internal system jargon), the window of exposure, what's being done for affected customers specifically, and what changed technically (again in plain terms: "we've added an additional access control layer and are auditing all similar configurations") without naming the specific internal service or engineer. Legal reviews the draft specifically for regulatory disclosure requirements in relevant jurisdictions before it ships, and the technical team confirms every factual claim in the customer version traces back to something actually verified in the internal postmortem, not to speculation.
Trade-offs and pitfalls
The most common failure is either two extremes: an overly technical public statement that reads as evasive because it's incomprehensible, or an overly vague one that reads as evasive because it says nothing concrete. A second common failure is treating legal review as a final rubber-stamp rather than involving it early enough to shape what can honestly and safely be said, which under time pressure to communicate fast, teams sometimes skip.
As the Systems Administrator, create a new user 'alice' with UID 1500, primary group 'developers', and home directory /home/alice. Set an initial password, configure password aging so it expires periodically, and demonstrate how to lock and unlock the account. Explain where password aging information is stored on the system and how you would inspect it for an existing user.
Sample Answer
Direct answer
useradd with explicit -u, -g, and -d flags creates the account with exactly the UID (user ID), primary group, and home directory specified rather than whatever the system would pick by default; chage manages password aging afterward, stored in /etc/shadow (not /etc/passwd, which holds the account's other fields but never the password aging policy); and locking an account is a distinct, reversible operation from deleting or disabling it, done by prefixing the password hash rather than removing it.
Creating the account
groupadd -g 2000 developers
useradd -u 1500 -g developers -m -d /home/alice -s /bin/bash alice
passwd alice
-u 1500: the explicit UID requested.-g developers: setsdevelopersas alice's primary group (must already exist, created above with a matching explicit GID for consistency across hosts if that matters to your environment, though the GID number itself was not specified in the requirement).-m: create the home directory if it does not already exist, populated from/etc/skel(the skeleton directory whose contents,.bashrcand similar default dotfiles, are copied into every new home directory).-d /home/alice: the explicit home directory path.-s /bin/bash: login shell.passwd alice: sets the initial password interactively.
Password aging
chage -M 90 -W 7 alice
-M 90: maximum 90 days between password changes, after which the password expires. -W 7: warn the user starting 7 days before that expiration. chage -l alice displays the current policy for an existing user, which is the command to reach for when inspecting rather than setting:
Last password change : Sep 24, 2026
Password expires : Dec 23, 2026
Password inactive : never
Account expires : never
Minimum number of days between password change : 0
Maximum number of days between password change : 90
Number of days of warning before password expires : 7
Where this is actually stored
Password aging fields live in /etc/shadow, one line per user, colon separated: username, the hashed password, and then the aging fields chage manages (last change date, minimum days, maximum days, warning days, inactivity period, and expiration date), all stored as integers, mostly counts of days since the Unix epoch. /etc/passwd holds the account's UID, GID, home directory, and shell, but never the password hash or aging policy on any current Linux system, which is exactly why /etc/shadow exists as a separate, more tightly permissioned file (root-only readable) in the first place.
Locking and unlocking
usermod -L alice
passwd -S alice
usermod -L (equivalent to passwd -l) prepends a ! to the password hash field in /etc/shadow, which makes the stored hash fail to match any input, blocking password authentication without touching the hash itself, meaning it can be reversed cleanly. passwd -S alice reports the current status compactly (L for locked, P for a usable password set, NP for no password set at all). Unlocking:
usermod -U alice
Worked example (executed)
groupadd -g 2000 developers
useradd -u 1500 -g developers -m -d /home/alice -s /bin/bash alice
id alice
chage -M 90 -W 7 alice
usermod -L alice
passwd -S alice
Real captured output:
uid=1500(alice) gid=2000(developers) groups=2000(developers)
alice L 2026-09-24 0 90 7 -1
L confirms locked status, and the fields after it (2026-09-24 0 90 7 -1) are exactly the last-change date, minimum, maximum, and warning days set above, plus an unset inactivity period (-1), read directly out of /etc/shadow by passwd -S, not a separate source of truth. Attempting usermod -U immediately afterward, before ever setting a real password with passwd, correctly refused with "unlocking the user's password would result in a passwordless account," a real safety check worth knowing about since it means unlock only works cleanly once an actual password hash exists to unlock back to.
Additional provisioning detail worth having ready
Beyond the base flow above: /etc/skel controls what a new home directory starts with, so a custom .bashrc or default config for new hires belongs there, not hand-copied per account. -G (capital, distinct from -g) adds supplementary group memberships beyond the primary group, for example useradd ... -G sudo,docker alice for a user who needs more than just her primary group's access. Forcing a password change at first login, common for a newly provisioned account whose initial password was set by an admin rather than chosen by the user: chage -d 0 alice, which sets the last-change date to the epoch, so the aging policy immediately considers the current password expired and prompts for a change on next login. To pair this provisioning flow with an automatic lockout after repeated bad login attempts (rather than only the manual usermod -L shown above), the same pam_faillock.so module used for account security policy elsewhere on the system can be configured with deny=3 in /etc/security/faillock.conf, so any account, alice's included, locks itself out after three consecutive failures without an administrator needing to intervene by hand each time.
Trade-offs and pitfalls
The most common mistake is locking an account with usermod -L and assuming that also blocks SSH key based login: it only disables password authentication, an account locked this way can still log in over SSH with a valid key if one is authorized, so a genuine "disable this account" action for a departing employee needs usermod -L plus either removing or invalidating their authorized_keys, and separately expiring the account itself (usermod -e 1 sets an account expiration date in the past, which does block all login methods, not just password). Second, chage -M alone with no -W warning gives a user zero notice before their password simply stops working, which reads as a mysterious lockout rather than an expected policy; always pair a maximum age with a reasonable warning window. Third, remember -g (primary group) and -G (supplementary groups) are genuinely different flags with different effects: confusing them either leaves a user without the intended primary group ownership on files they create, or fails to grant a supplementary group's access at all.
How do group types and scopes work in a multi-domain forest, and how would you nest groups to grant access to a resource? What risk comes from converting a distribution group to a security group?
Sample Answer
Direct answer
A group has a type and a scope. (Terms used below: a SID, security identifier, is the unique value that identifies an account or group inside permission lists and access tokens; a trusting domain or forest is one that has agreed to let accounts from another domain or forest be granted access to its resources.) The type says what it is for: a security group can be listed in access control lists (DACLs, the permission lists on files, shares and objects), a distribution group can only be used for email. The scope says who can be a member and where the group can be granted access. In a multi-domain forest use AGDLP: Accounts go into Global groups (one per role, in the user's own domain), global groups are nested into Domain Local groups (one per resource and permission level, in the resource's domain), and permissions are granted to the domain local group. Converting a distribution group to a security group is risky because the list's membership, curated for email, suddenly becomes a source of access.
Scopes and membership rules
Why the scopes sit where they do: a global group only holds members from its own domain, so its membership changes replicate inside that domain and nowhere else. A universal group's membership is stored in the global catalog and replicates forest-wide, so it can hold members from anywhere but every change costs forest-wide replication. A domain local group is visible for permissions only in its own domain, which is why it belongs next to the resource.
| Scope | Members can be | Can be used for permissions | Can be converted |
|---|---|---|---|
| Global | Accounts and global groups from the same domain | Anywhere in the forest, and in trusting domains and forests | To universal, if it is not a member of another global group |
| Domain Local | Accounts, global groups and universal groups from any domain in the forest or a trusted domain; domain local groups from the same domain | Only in its own domain | To universal, if it contains no other domain local group |
| Universal | Accounts, global groups and universal groups from any domain in the forest | Anywhere in the forest or in trusting forests | To domain local, if it is not a member of another universal group; to global, if it contains no other universal group |
The conversion column follows from the nesting rules. A universal group cannot be a member of a global group, so a global group that is a member of another global group cannot become universal. Example: GG-Finance-EU sits inside GG-All-EU, so converting GG-Finance-EU to universal fails until it is removed from GG-All-EU. In the same way, a domain local group holding another domain local group cannot become universal, and a universal group holding other universal groups cannot become global.
Local groups on a member server or workstation are a fourth kind of group, stored on that computer. Per Microsoft's table, a domain local group can be a member of local groups on computers in the same domain (except built-in groups with well-known SIDs). Use local groups for per-machine rights, not for shared resources.
Nesting for a resource (worked example)
Forest: corp.contoso.com (parent) and eu.contoso.com (child domain). The share \EUFS01\Finance sits in the eu domain. Finance staff exist in both domains.
- Create global group GG-Finance-Corp in corp and GG-Finance-EU in eu. Each holds only its own domain's finance users.
- Create domain local group DL-EUFS01-Finance-Modify in eu and put both global groups in it (a domain local group accepts global groups from any domain).
- Grant Modify on the Finance share and NTFS folder to DL-EUFS01-Finance-Modify only.
- A new finance hire is added to the global group in their own domain. Nothing on the file server changes.
When one role is needed in many resource domains, nest the global groups into a universal group (UG-Finance-All) and nest that into the domain local groups. Keep the universal group's direct members few and stable, because a domain controller must be able to supply universal group membership at logon from a global catalog or a cached copy. Link-value replication (LVR) replicates only the one added or removed member value instead of the whole member list, which matters for groups with thousands of members. It needs a forest functional level of Windows Server 2003 or higher (the functional level is the setting that says which Windows Server versions the domains and forest must run, and it unlocks features accordingly).
Do not put users directly into the domain local group: that works until the second domain or the second permission level arrives, and then every ACL (access control list) needs re-editing. Permissions also come in two layers on a file share, share permissions and NTFS permissions on the folder; a common pattern is to open the share to the same group and let NTFS do the fine-grained control, since the more restrictive of the two wins.
Token size
Every group SID (security identifier) a user belongs to goes into that user's access token at logon. The flattened set means the token lists every group the user is in, including the ones reached only through nesting, so nesting does not reduce the count. There are two separate limits. The Windows access token holds at most about 1,010 group SIDs (Microsoft documents a 1,010 SID limit for the LSA access token). Kerberos tickets carry the same data in a size-limited field, and since Windows Server 2012 the default MaxTokenSize is 48,000 bytes. Microsoft's estimate of the ticket size is 1200 + 40d + 8s bytes, where d counts universal-group memberships from other domains plus SID history entries, and s counts memberships in global groups, domain local groups and universal groups from the user's own domain.
Worked check with illustrative numbers: a user in 600 global groups in their own domain has s = 600, so the ticket is about 1200 + 8 x 600 = 6,000 bytes, far under 48,000. Solving the formula for the limit, 48,000 allows (48,000 - 1,200) / 8 = 5,850 such groups, or (48,000 - 1,200) / 40 = 1,170 cross-domain universal memberships. For ordinary global-group memberships the 1,010 SID limit is reached first, at 1200 + 8 x 1,010 = 9,280 bytes, long before the byte limit. Symptoms of token bloat are logon failures or access denied for users in many groups. Rationalize group membership rather than only raising the limit.
Converting a distribution group to a security group
A distribution group is not security-enabled, so it cannot be in a DACL. Conversion makes it a security principal, and its SID goes into the token of every member from their next logon. The risks:
- The membership was maintained for mail, not for access: it can include stale members, broad nested lists and people who should never see a given resource. Anything that later grants this group access, or already references it, now applies to all of them.
- Token growth: every converted group adds a SID to each member's token.
- The conversion is silent. Nothing alerts the owner of the file share that a mailing list is now a permission source.
Safer path: audit and clean the membership first, or create a new security group with the intended members instead of converting, and review any place the group name already appears in permissions.
Pitfalls
- Putting accounts straight into ACLs instead of groups.
- Using universal groups everywhere. They make logon depend on global catalog access and increase what must be available forest-wide.
- Nesting too deeply: it hides who has access and makes audits slow.
How would you build an automated pipeline that continuously proves your backups are both technically valid and actually restorable, without a human manually running a restore every time? Sketch how it would work end-to-end and what would make you trust its result.
Sample Answer
Direct answer
An automated pipeline that continuously proves backups are both technically valid and actually restorable works by generating known test data, backing it up through the real backup path, restoring it into an isolated environment through the real restore path, and verifying the restored data cryptographically matches the original, all on a schedule, with the trust coming specifically from the fact that it exercises the same restore mechanism a real recovery would use rather than only checking that a backup file exists.
Structured elaboration
End-to-end sketch. On a schedule (for example, nightly per system, or continuously in small batches across a large fleet), the pipeline generates a fresh, deterministic block of synthetic canary data, seeds it into the source system being protected, waits for that data to be captured by the normal backup process (or triggers an out-of-band backup of just that canary data), then restores it, using the real restore mechanism, into a disposable, isolated target rather than production. It then compares the restored data against the original canary data byte-for-byte via a cryptographic hash, and reports a clear pass or fail signal, along with how long each phase took, to a dashboard and an alerting system.
What would make you trust its result. Trust comes from the pipeline using the same code paths a real disaster recovery would use, the actual backup job and the actual restore mechanism, not a shortcut or a simulation of them; from verifying at the content level (a hash of the restored bytes matching the original) rather than only checking that a restore operation returned a success status code, since a restore can report success while silently delivering wrong or incomplete data; from running frequently and automatically enough that a regression is caught within a day or two rather than months later during a real incident; and from the canary data being deterministic and independently reproducible, so a failure can be re-run and confirmed rather than dismissed as a one-off glitch.
Worked example
The block below is a minimal, runnable simulation of exactly this shape: it generates a synthetic dataset with a fixed random seed (so the "before" state is reproducible), computes a checksum of every file, runs it through a stand-in backup step and a stand-in restore step, computes checksums of the restored files, and reports whether every file came back byte-identical. In a real system, the backup and restore functions would call the organization's actual backup and restore APIs (a cloud backup service, a database client, an agent) instead of a local file copy, and the pipeline would run this on a schedule against real systems rather than synthetic local files, but the verification logic, generate known data, capture its checksums, back it up, restore it, and compare checksums on the other side, is the same shape regardless of scale.
import hashlib
import os
import random
import shutil
import tempfile
import time
def sha256_of_file(path):
h = hashlib.sha256()
with open(path, "rb") as f:
for chunk in iter(lambda: f.read(65536), b""):
h.update(chunk)
return h.hexdigest()
def generate_synthetic_dataset(target_dir, num_files=20, seed=42):
rng = random.Random(seed)
os.makedirs(target_dir, exist_ok=True)
for i in range(num_files):
path = os.path.join(target_dir, f"file_{i:03d}.dat")
size_bytes = rng.randint(1024, 8192)
with open(path, "wb") as f:
f.write(rng.randbytes(size_bytes))
return target_dir
def snapshot_checksums(root_dir):
checksums = {}
for name in sorted(os.listdir(root_dir)):
path = os.path.join(root_dir, name)
checksums[name] = sha256_of_file(path)
return checksums
def backup(source_dir, backup_dir):
# Stand-in for the real backup API call.
if os.path.exists(backup_dir):
shutil.rmtree(backup_dir)
shutil.copytree(source_dir, backup_dir)
def restore(backup_dir, restore_dir):
# Stand-in for the real restore API call.
if os.path.exists(restore_dir):
shutil.rmtree(restore_dir)
shutil.copytree(backup_dir, restore_dir)
def run_restore_drill(num_files=20, seed=42):
workdir = tempfile.mkdtemp(prefix="restore_drill_")
source_dir = os.path.join(workdir, "source")
backup_dir = os.path.join(workdir, "backup")
restore_dir = os.path.join(workdir, "restore")
start = time.time()
generate_synthetic_dataset(source_dir, num_files=num_files, seed=seed)
source_checksums = snapshot_checksums(source_dir)
backup(source_dir, backup_dir)
restore(backup_dir, restore_dir)
restored_checksums = snapshot_checksums(restore_dir)
elapsed = time.time() - start
mismatches = [
name for name in source_checksums
if source_checksums.get(name) != restored_checksums.get(name)
]
missing = sorted(set(source_checksums) - set(restored_checksums))
passed = not mismatches and not missing
print(f"files_checked={len(source_checksums)}")
print(f"mismatched_files={len(mismatches)}")
print(f"missing_files={len(missing)}")
print(f"elapsed_seconds={elapsed:.4f}")
print(f"result={'PASS' if passed else 'FAIL'}")
shutil.rmtree(workdir)
return passed
if __name__ == "__main__":
run_restore_drill(num_files=20, seed=42)
Running this script unchanged in a fresh directory with only the Python standard library available prints output like:
files_checked=20
mismatched_files=0
missing_files=0
elapsed_seconds=0.0637
result=PASS
The first four values are deterministic given the fixed seed of 42, files_checked will always read 20, mismatched_files and missing_files will always read 0, and result will always read PASS, for this unmodified backup and restore logic; elapsed_seconds is a wall-clock timing measurement and will vary slightly by machine and disk speed, but the pass/fail verdict itself does not depend on timing.
Trade-offs and pitfalls
- Verifying only that the restore operation returned success, without independently re-checking content, misses exactly the failure class that matters most: a restore that completes without error but silently delivers wrong or stale bytes. The hash comparison in the example above is what closes that gap, and a real pipeline should keep that same principle even when the backup and restore steps are swapped for real production APIs.
- Running this constantly against synthetic canary data proves the backup and restore mechanism itself works, but it does not by itself prove every real dataset in production is being backed up correctly (a canary passing says nothing about a completely different, unmonitored system); it needs to be paired with per-system coverage tracking, not treated as a universal proxy for "everything is fine."
- A failure in the isolated test environment itself (out of disk space, a misconfigured target) can look identical to a genuine backup or restore failure; a production version of this pipeline should retry once in a fresh environment before treating a failure as a confirmed, alertable incident, to avoid false alarms from the test harness rather than the thing it's testing.
Recommended Additional Resources
- The Linux Command Line by William Shotts (free online) - comprehensive Linux fundamentals
- Windows Server 2019 Administration Fundamentals (Microsoft Learn)
- Kubernetes in Action by Marko Luksa - for understanding container orchestration if relevant to role
- System Design Primer (GitHub) - infrastructure design patterns and principles
- High Performance MySQL by Baron Schwartz - for database administration and backup strategies
- DevOps Handbook - for understanding infrastructure automation and operational excellence
- Incident Response and Recovery by Andrew Cormack - practical disaster recovery and incident management
- Linux Academy / A Cloud Guru / Pluralsight - hands-on Linux and Windows administration labs
- AWS/Azure/GCP free tier documentation - understand major cloud platforms if company uses them
- Cracking the Systems Design Interview (GitHub repositories) - interview preparation for infrastructure roles
- Practice platforms: Hack The Box, TryHackMe (systems administration challenges), Linux Journey (interactive Linux learning)
- Study FAANG company blog posts on infrastructure: Google Cloud Architecture patterns, AWS Well-Architected Framework, Microsoft Azure Architecture
- Prepare using: mock interview platforms (Pramp, Interviewing.io), LeetCode (system design discussions), structured interview notes
Search Results
Operating System Interview Questions - GeeksforGeeks
Intermediate OS Interview Questions · 61. Write a difference between a user-level thread and a kernel-level thread? · 62. Write down the advantages of ...
Fundamentals Linux MCQs for System Administrators
These questions test knowledge on the core components of the Linux operating system, such as the kernel, and essential directories.
10 Common SAP Basis Interview Questions (With Answers) - Indeed
1. How is SAP used in organisations? · 2. What are the responsibilities and duties of an SAP Basis administrator? · 3. What are SAP instances and why do we create ...
10 Killer Questions That Flip SysAdmin Interviews Upside Down
Preparing for a Systems Administrator interview and want to stand out? In this video, we reveal the essential questions you should ask the interview panel ...
50 Most Popular Salesforce Interview Questions & Answers ...
1. Describe how Salesforce CRM is used by organizations? At its core, Salesforce is a customer-facing CRM system. It is used to record customer ...
▷ Top 40+ Azure Interview Questions and Answers - igmGuru
2. For Azure, what does a role instance mean? 3. How does Azure Diagnostics API help organizations? 4. Explain different cloud deployment models. 5. What is the ...
90+ AWS Interview Questions and Expert Answers (2025)
Most Asked AWS interview questions and expert answers for freshers to experienced professionals. Master EC2, S3, Lambda & more to crack AWS job interview in ...
Top 110+ DevOps Interview Questions and Answers for 2026
Here are some of the most common DevOps interview questions and answers that can help you while you prepare for DevOps roles in the industry.
This interview preparation guide was generated using AI-powered research from the sources listed above. While we strive for accuracy, we recommend verifying critical information from official company sources.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Systems Administrator jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs