DoorDash Systems Administrator (Junior Level) - Interview Preparation Guide
DoorDash's interview process for Systems Administrator positions typically follows a structured format designed to assess technical fundamentals, hands-on infrastructure knowledge, problem-solving abilities, and cultural fit. For a junior-level candidate, the process emphasizes foundational knowledge of operating systems, networking, server management, and the ability to learn and grow in the role. Expect a combination of phone-based technical screening and onsite interviews with multiple team members.
Interview Rounds
Recruiter Screening
What to Expect
This is a brief initial screen with a technical recruiter to assess your background, interest in the role, and basic qualifications. The recruiter will confirm your technical background, understanding of the role, and general fit. This is also an opportunity to ask questions about the team, company culture, and next steps. Typically 20-30 minutes.
Tips & Advice
Be enthusiastic about infrastructure and systems administration. Clearly articulate your understanding of what a Systems Administrator does. Have questions ready about the team size, their infrastructure stack, and growth opportunities. This round is mostly about communication skills and cultural alignment, not technical depth.
Focus Topics
Communication Skills
Clearly explaining technical concepts in understandable terms without jargon
Practice Interview
Study Questions
Background and Experience Summary
Concisely explaining your technical background, any internships, personal projects, or relevant coursework
Practice Interview
Study Questions
Role Understanding and Motivation
Demonstrating clear understanding of what Systems Administrators do and why you're interested in the role
Practice Interview
Study Questions
Technical Phone Screen
What to Expect
A technical interview conducted over phone or video with an engineer from the infrastructure or systems team. This round focuses on foundational knowledge of operating systems, server administration basics, networking fundamentals, and troubleshooting approach. You'll be asked scenario-based questions about how you'd handle common infrastructure tasks and problems. Expect questions about Linux commands, Windows administration, user management, backups, and basic network concepts. The interviewer is assessing your technical foundation and how you approach problem-solving. Approximately 45-60 minutes.
Tips & Advice
Review Linux command-line fundamentals extensively (file permissions, user management, package managers, process management, logs). Be comfortable with Windows Server basics as well. When answering scenario questions, walk through your thought process step-by-step rather than jumping to conclusions. For junior candidates, the interviewer expects solid foundational knowledge but understands you won't know everything—it's okay to say 'I'm not sure, but here's how I'd find out.' Practice explaining your reasoning for troubleshooting decisions. Have real or hypothetical examples ready from your experience or projects.
Focus Topics
Networking Fundamentals
OSI model, TCP/IP, DNS, DHCP, IP addressing, subnetting basics, routing concepts, and firewall basics
Practice Interview
Study Questions
Backup and Disaster Recovery Concepts
Basic backup strategies, retention policies, recovery procedures, and the importance of verification
Practice Interview
Study Questions
Windows Server Administration Basics
Server roles, Active Directory basics, Group Policy, user account management, file sharing, Windows services, and event logs
Practice Interview
Study Questions
Linux Operating System Fundamentals
Command-line usage, file systems, permissions (chmod, chown), user and group management, package management, process management, and log files
Practice Interview
Study Questions
Troubleshooting Methodology
Systematic approach to identifying and solving problems: gathering information, forming hypotheses, testing solutions, and documenting results
Practice Interview
Study Questions
User Account and Access Management
Creating and managing user accounts, setting permissions, sudo/elevation, password policies, and access control principles
Practice Interview
Study Questions
Onsite Round 1 - Infrastructure and System Administration Fundamentals
What to Expect
First onsite interview with a senior Systems Administrator or infrastructure engineer. This round dives deeper into practical system administration skills, including server configuration, operating system knowledge, and hands-on infrastructure tasks. You may be asked to diagram network architectures, explain configuration decisions, or discuss how you'd approach setting up systems. The interviewer evaluates your technical depth, understanding of infrastructure best practices, and ability to ask clarifying questions. Approximately 45-60 minutes.
Tips & Advice
Be prepared to discuss real-world infrastructure scenarios and how you'd approach them. For example, you might be asked 'How would you set up a new user with appropriate access across multiple systems?' or 'Walk me through how you'd diagnose a slow server.' Show your work and explain your reasoning. It's better to ask clarifying questions (e.g., 'What's the current infrastructure like?' or 'What are the security requirements?') than to make assumptions. Draw diagrams if it helps explain your thinking. Relate answers back to the job description responsibilities where possible.
Focus Topics
Storage and File Systems
Understanding different file systems (ext4, NTFS), managing disk partitions, quotas, and storage architecture
Practice Interview
Study Questions
Disaster Recovery Planning
Creating recovery plans, RTO/RPO concepts, failover strategies, and documenting recovery procedures
Practice Interview
Study Questions
Software Installation and Patching
Installing applications, managing software licenses, applying security patches and updates, managing dependencies
Practice Interview
Study Questions
Security Fundamentals for Infrastructure
Access control, encryption basics, firewall configuration, intrusion detection, security patching, and security best practices
Practice Interview
Study Questions
Server Installation and Configuration
Setting up and configuring servers from scratch, including OS installation, hardware configuration, drivers, and initial security setup
Practice Interview
Study Questions
System Monitoring and Performance Management
Tools and techniques for monitoring CPU, memory, disk usage, network utilization; understanding performance baselines and alerts
Practice Interview
Study Questions
Onsite Round 2 - Networking and Infrastructure Troubleshooting
What to Expect
Second onsite interview with a network engineer or infrastructure specialist from the team. This round focuses on networking knowledge, network infrastructure management, and practical troubleshooting scenarios. You may be presented with real-world network problems and asked to diagnose and solve them. Topics include network design concepts, routing, switching, DNS/DHCP, VLANs, and common infrastructure issues. The interviewer assesses your ability to think systematically about network problems and your practical networking knowledge. Approximately 45-60 minutes.
Tips & Advice
Practice troubleshooting network problems methodically. For example, if asked about connectivity issues, walk through checking DNS, DHCP, routing, firewall rules, and physical connectivity in a logical order. Use tools like ping, traceroute, netstat, and ifconfig/ipconfig conceptually. Understand the OSI model and be able to map problems to layers. If you're presented with a network diagram, ask questions before diving in. For a junior role, deep networking expertise isn't expected, but showing a systematic troubleshooting approach is essential. Have examples ready of times you diagnosed networking issues.
Focus Topics
Remote Access and VPN Basics
VPN concepts, remote access technologies, SSH tunneling, and secure remote administration
Practice Interview
Study Questions
Routing and Switching Concepts
Understanding routing protocols, switch configuration, VLANs, spanning tree, and how traffic flows through networks
Practice Interview
Study Questions
Network Architecture and Design Basics
Understanding network topologies, segmentation, VLANs, subnetting, and how to design networks for scalability and security
Practice Interview
Study Questions
Firewall and Security Infrastructure
Firewall rules, access control lists, network security policies, and how security controls are implemented in infrastructure
Practice Interview
Study Questions
DNS and DHCP Configuration
How DNS and DHCP work, configuring these services, troubleshooting DNS resolution and IP assignment issues
Practice Interview
Study Questions
Network Troubleshooting and Diagnostics
Systematic approach to diagnosing connectivity issues, using diagnostic tools, understanding network protocols and common failure modes
Practice Interview
Study Questions
Onsite Round 3 - Behavioral and Culture Fit
What to Expect
Final onsite interview with a team lead, manager, or senior team member focused on behavioral assessment and cultural alignment. This round evaluates how you work in teams, handle challenges, approach learning and growth, and align with company values. You'll be asked behavioral questions about past experiences, how you handle pressure, collaboration, and learning from mistakes. The interviewer also assesses your enthusiasm for the role and fit with the team culture. Approximately 45-60 minutes.
Tips & Advice
Prepare specific examples using the STAR method (Situation, Task, Action, Result) for common behavioral questions like 'Tell me about a time you made a mistake and how you handled it' or 'Describe a challenging technical problem you solved.' For a junior role, emphasize your learning ability, willingness to take on new challenges, collaboration with team members, and how you handle not knowing something. Be genuine about what you don't know and your eagerness to learn. Ask thoughtful questions about the team's projects, culture, and how they approach mentoring junior team members. Show interest in the company's infrastructure and operations.
Focus Topics
Handling Failure and Mistakes
How you respond to making mistakes, learning from errors, and preventing similar issues in the future
Practice Interview
Study Questions
Adaptability and Handling Pressure
How you handle changing priorities, on-call situations, and stressful infrastructure incidents with calm and focus
Practice Interview
Study Questions
Technical Communication
Explaining technical concepts to non-technical stakeholders, documenting procedures, and asking clarifying questions
Practice Interview
Study Questions
Initiative and Ownership
Taking responsibility for tasks, proactively identifying improvements, and not waiting to be told what to do
Practice Interview
Study Questions
Teamwork and Collaboration
Working effectively with team members, communicating technical information, receiving feedback, and contributing to team goals
Practice Interview
Study Questions
Problem-Solving and Learning Approach
How you approach unfamiliar problems, your willingness to learn, resourcefulness in finding answers, and persistence in troubleshooting
Practice Interview
Study Questions
Frequently Asked Systems Administrator Interview Questions
Write a cloud-init user-data YAML snippet for an Ubuntu instance that creates a user named 'deploy', uploads a provided SSH public key, updates apt package lists, installs nginx and a monitoring agent, and ensures the nginx service is started. Provide a clear, minimal cloud-init example in YAML.
Sample Answer
Approach
- Create a user "deploy" with sudo and an SSH public key.
- Ensure apt lists are updated and required packages installed.
- Start and enable nginx; install a monitoring agent (example: datadog-agent placeholder).
- Keep the snippet minimal and ready to paste as cloud-init user-data.
cloud-init user-data (YAML)
#cloud-config
users:
- name: deploy
gecos: Deploy User
sudo: ALL=(ALL) NOPASSWD:ALL
shell: /bin/bash
ssh_authorized_keys:
- "ssh-rsa AAAA...your-public-key... user@example.com"
package_update: true
packages:
- nginx
- curl
- apt-transport-https
runcmd:
- [ sh, -c, "systemctl enable --now nginx" ]
- [ sh, -c, "echo 'Installing monitoring agent (example)'; curl -sS https://example.com/install-agent.sh | bash" ]
Notes:
- Replace the SSH key and monitoring-agent install command/URL with your actual public key and vendor instructions (e.g., Datadog, Prometheus node-exporter).
- For production, pin package versions or add apt repositories and GPG keys securely.
Define progressive disclosure and describe two concrete ways you would use it in technical documentation so a reader can go from a high-level decision down to low-level implementation detail without being overloaded.
Sample Answer
Direct answer
Progressive disclosure means showing the minimum someone needs to make their next decision first, then letting them opt into more detail only if they need it, rather than presenting every layer of a decision at once. For technical documentation, that means separating what we decided and why it matters from how it's actually implemented, and only showing the second layer to someone who asks for it.
Structured elaboration
Two concrete ways to build this into documentation:
- A "decision, then detail" page structure. The top of the page states the decision and its business-relevant effect in one or two sentences. Directly below, an expandable or clearly linked section holds the reasoning (why this option over the alternatives, what constraint drove it), and a separate section holds the implementation (exact commands, config, code). A reader making a go/no-go call never has to scroll past architecture detail to find the decision.
- Collapsed detail blocks inside a page that stays otherwise readable. Long code blocks, diagrams, or benchmark tables default to collapsed, with a label that tells the reader what's inside before they open it, not just "details." This keeps the page skimmable top to bottom for someone doing a first pass, while an engineer implementing the change can expand everything in order.
Both work because they let the reader choose their own depth instead of the writer choosing it for them, and because the label on each layer, decision, why, how, tells the reader which layer they're in before they commit to reading it.
Worked example
A raw engineering note might read: "Switched to regional read replicas with async replication and connection pooling via PgBouncer to cut p95 read latency." Applied with progressive disclosure, the page becomes:
Top line (decision layer): "We added copies of the database closer to users in each region so read requests don't have to cross the country, which is what was making some pages feel slow for customers far from our main data center."
Expandable "why" layer: explains the latency problem was concentrated in specific regions and why a cache alone wasn't sufficient, still in plain language.
Expandable "implementation" layer, collapsed by default: regional read replicas (copies of the database kept near each user region), updated by async replication (the copy is written a short delay after the original, not instantly), and connection pooling via PgBouncer (a tool that reuses open database connections instead of opening a new one per request), plus config snippets and the failover procedure.
A reader deciding whether to approve the change never has to parse "PgBouncer" or "async replication" to get the decision; an engineer implementing it clicks straight through to exactly that.
Trade-offs and pitfalls
Progressive disclosure can misfire if the top layer is vague instead of just simple: "we improved performance" tells the reader nothing they can act on, while "reads are faster for users far from our main region" does. It also fails if the label on a collapsed section doesn't say what's inside; readers won't expand something called "details," so label it with what they'll actually get, for example "config and rollback steps." And it isn't free: every layer you maintain is another thing that can drift out of sync with the code, so it's worth it for docs people repeatedly return to, not a one-off internal note nobody will reread.
You encounter a kernel panic on a Linux production host or a Blue Screen on Windows. Describe immediate triage steps you would take to gather information, preserve evidence, minimize downtime, and safely restore service. Include commands and tools you would run if the system is still reachable over the network.
Sample Answer
Immediate goals: preserve evidence, gather diagnostic data, minimize downtime (failover/rollback), avoid further writes to damaged system.
Initial triage (both OSes)
- Verify reachability and alerts, note time of crash and affected services.
- Put host into maintenance mode (disable monitoring/automated remediation) and if possible shift traffic to standby.
If system still reachable (Linux)
- Capture uptime/last boot and kernel messages:
who -b
journalctl -k -b -1 # logs from previous boot
dmesg -T | tail -n 200
- Preserve crash dump / vmcore:
# check kdump
systemctl status kdump
ls -l /var/crash /var/lib/kdump
# copy vmcore off-box
scp /var/crash/<vmcore> analyst@collector:/data/
- Trigger safe sysrq (if hung and acceptable):
echo s > /proc/sysrq-trigger # sync
echo u > /proc/sysrq-trigger # remount ro
echo b > /proc/sysrq-trigger # reboot (only if necessary)
- Collect config and state:
tar czf /tmp/sysinfo.tgz /etc /var/log
ss -tunap > /tmp/conns.txt
ps aux > /tmp/ps.txt
scp /tmp/sysinfo.tgz analyst@collector:/data/
- If memory dump missing, use makedumpfile/crash utilities to extract vmcore.
If system still reachable (Windows BSOD remote collection)
- Retrieve minidump and MEMORY.DMP:
- Copy from \host\C$\Windows\Minidump and C:\Windows\MEMORY.DMP to safe storage.
- Query Windows event logs remotely:
Get-WinEvent -ComputerName host -FilterHashtable @{LogName='System';ID=1001} | Format-List
wevtutil qe System /q:"*[System[(EventID=1001)]]" /f:text > system_bsod.txt
- Check crash settings:
Get-WmiObject -Class Win32_ComputerSystem -ComputerName host
- If machine down, collect SAN snapshots or disk images rather than booting to avoid altering evidence.
Forensics & analysis
- Preserve timestamps and chain-of-custody (who copied which files and when).
- Analyze vmcore/minidump with crash, makedumpfile, WinDbg:
- Linux: crash /usr/lib/debug/lib/modules/... /path/to/vmcore
- Windows: WinDbg -z MEMORY.DMP !analyze -v
Restore service safely
- If quick rollback possible, failover to replica or boot last known-good image.
- If repair required, boot into rescue mode, apply kernel rollback or driver updates, test in staging, then bring back gradually.
- Postmortem: root cause, mitigations (kdump enabled, kernel/drivers pinned, monitoring alerts), and update runbooks.
Key priorities: copy crash dumps off-box, avoid writing to evidence disks, restore service via failover, and perform controlled analysis offline.
Leadership/case study: You have a backlog of long-running performance engineering projects and a queue of production incidents demanding immediate attention. Propose a prioritization framework that balances short-term reliability fixes, long-term performance investments, and feature delivery. Include metrics you would track to demonstrate ROI of performance work and how you would get leadership buy-in.
Sample Answer
Overview / goal
I’d balance immediate reliability, long-term performance, and new features by using a transparent, score-based prioritization that ties technical work to business impact and SLAs.
Prioritization framework
- Score work by: Impact (customer-facing SLA or revenue risk), Urgency (incident vs. planned), Effort (person-days), Risk reduction (likelihood of future incidents).
- Use WSJF-like formula: Priority = (Impact + Urgency + RiskReduction) / Effort.
- Reserve a weekly incident SLA lane (e.g., >30% capacity) for fire-fighting; allocate remaining capacity to a mix: 50% reliability/perf projects, 30% feature ops, 20% innovation/tech debt.
- Example items: patching critical CVE (high urgency, low effort), DB query tuning (high impact, medium effort), instance right-sizing (cost-saving, low risk).
Metrics to demonstrate ROI
- Operational: MTTR, MTBF, incident count by severity, on-call hours.
- Performance: P95 / P99 latency, error rate, throughput.
- Financial: cloud spend before/after (e.g., $/req), avoided outage revenue loss.
- Business: % of SLAs met, customer support tickets related to performance.
Getting leadership buy-in
- Present prioritized roadmap with quantified impact and cost estimates, plus a 90-day pilot showing measurable wins (e.g., reduced P95 by X ms, saved $Y/month).
- Tie improvements to business KPIs (revenue, retention, SLA penalties).
- Provide executive dashboard and monthly playbooks showing trends and risk posture.
- Offer a “stop-the-bleeding” SLA guarantee: incidents addressed within defined window to reassure stakeholders.
System design task: You manage a small web application with three web servers behind a load balancer, a PostgreSQL database, and a 500 GB NFS file share containing user uploads. Business requires RTO <= 4 hours and RPO <= 1 hour. Propose a backup and recovery architecture (backup types, schedules, retention, storage targets, and test plans) that meets the requirements and minimizes cost.
Sample Answer
Situation & objectives
RTO ≤ 4 hours, RPO ≤ 1 hour. Small app: 3 web servers, Postgres DB, 500 GB NFS user uploads. Minimize cost while meeting SLAs.
Backup strategy (high-level)
- PostgreSQL: continuous WAL shipping (streaming replication to a standby or archive WALs) + nightly base backups (pg_basebackup).
- NFS uploads: hourly incremental snapshots (filesystem-level or block-level) + daily full snapshot.
- Web servers: stateless; keep config as code (Ansible/Git). Back up configs + package lists nightly.
Schedules & RPO/RTO mapping
- WAL streaming enables point-in-time recovery to ≤ 1 hour RPO (WAL upload every 1–5 min).
- Base backup: nightly full (keep 7 days). WAL retention to cover at least 24–48 hours.
- NFS snapshots: hourly retention 24 hours, daily full retention 14 days, weekly archives 90 days (move older to colder storage).
Storage targets & cost
- Primary backups to cheap object storage (S3/Glacier or equivalent). Keep latest 48h snapshots on faster storage (S3 Standard/IA) for quick restore; archive older to Glacier/Coldline.
- Replicate critical DB WALs to a second region/cloud account if budget allows.
Recovery plan & RTO
- DB failover to standby (minutes) or restore from base + WAL replay: estimated < 2 hours.
- NFS restore: restore latest hourly snapshot to a new volume and remount; 500 GB from object-store snapshot to fast block storage (provision parallel workers) — target < 2 hours.
- Web servers redeploy from config repo → < 30 minutes.
Testing & verification
- Quarterly full DR test: simulate db loss and NFS loss; restore into isolated VPC, verify app end-to-end.
- Monthly partial tests: restore one hourly snapshot and apply WALs to validate integrity; verify checksums and file counts.
- Daily automated backup validation: verify snapshot creation, WAL archive success, and run quick checksum.
Retention & cleanup
- Hourly: keep 24
- Daily: keep 14
- Weekly: keep 12 (months)
- Monthly/Archive: keep 1 year on cold storage
Operational notes
- Automate with cron/backup tool (Borg/Restic for files + lifecycle policy to object storage); use monitoring and alerts for failed backups.
- Encrypt backups at rest and in transit; restrict access via IAM.
- Document runbooks for restore steps and maintain RTO timers.
Describe how you would use the Delegation of Control Wizard or AD ACLs to allow a Help Desk group to reset passwords and unlock accounts for users in the 'Employees' OU without granting additional rights. Outline steps to configure auditing to track these delegated actions and any procedural controls to ensure least privilege and accountability.
Sample Answer
Situation / Goal
I needed Help Desk to reset passwords and unlock accounts only for users in the Employees OU, with full auditability and least-privilege controls.
Delegation steps (practical)
- Open Active Directory Users and Computers, right-click Employees OU → Delegate Control.
- Add the Help Desk group.
- Use the wizard: choose “Create a custom task to delegate” → “Only the following objects in the folder” → select “User objects.”
- Grant permissions: “Reset user passwords and force password change at next logon.”
- If unlock requires explicit ACL change, use Advanced Security on the OU to grant Write lockoutTime (or use granular ACE via “Delegate Control” to add that right).
If using AD ACLs directly
- Open ADUC → View → Advanced Features → OU Properties → Security → Advanced → Add Help Desk group → select specific Extended Rights: “Reset Password” and allow write lockoutTime.
Auditing configuration
- Enable success auditing for “Account Management” in Group Policy (Computer Configuration → Policies → Windows Settings → Security Settings → Advanced Audit Policy Configuration → Account Management → Audit User Account Management: Success).
- Configure SACL on Employees OU (Advanced Security → Auditing) to audit Success for Reset Password and Write lockoutTime.
- Forward logs to SIEM; alert on Event IDs: 4724 (password reset attempt), 4767 (account unlocked), and correlate with ticket IDs.
Procedural controls for least privilege & accountability
- Require ticket with unique ID and approval before Help Desk action; log ticket ID in AD operation notes.
- Enforce MFA for Help Desk consoles and privileged sessions.
- Monthly access review (recertification) of the Help Desk group memberships.
- Implement Just-In-Time elevation (Privileged Access Management) for sensitive cases.
- Retain audit logs 1 year; run weekly automated reports of resets/unlocks and investigate anomalies.
Result: Help Desk can perform necessary support without expanded rights, all actions are auditable and controlled, preserving least privilege and accountability.
What's the difference between redundancy and replication when it comes to service reliability? Walk through an example for a stateless service and a stateful service, and name a failure mode that redundancy alone doesn't protect against for the stateful one.
Sample Answer
Direct answer: Redundancy is having extra, interchangeable components standing by so one can take over if another fails; replication is actively keeping copies of state (data) synchronized across multiple nodes so the state itself survives a failure, not just the compute that serves it. They're often used together, but they solve different problems: redundancy alone is enough for a stateless service, because any interchangeable instance can serve any request; a stateful service needs replication too, because a fresh redundant instance with no data isn't actually a working replacement.
Structured elaboration
| Aspect | Redundancy (stateless) | Replication (stateful) |
|---|---|---|
| What's duplicated | Compute/serving capacity | Data/state itself |
| Failover requirement | Route traffic to a healthy instance; done | Promote a replica that has the data, and ensure it's sufficiently up to date |
| Consistency concern | None, any instance is interchangeable | Central concern: how in-sync are the replicas at failover time |
| Typical mechanism | Load balancer + auto-healing instance group | Leader-follower or multi-leader data replication |
Stateless example: a set of identical app server instances behind a load balancer, handling API requests with no local state. If one instance dies, the load balancer routes around it and an autoscaler replaces it; the new instance needs no data transfer because there was never any instance-local state to lose. This is pure redundancy: extra interchangeable copies of the same stateless computation.
Stateful example: a primary-replica database. The primary accepts writes; replicas continuously receive a copy of the write stream (replication). If the primary fails, a replica is promoted to take over. Unlike the stateless case, simply having an extra database instance running (redundancy alone, no replication) would give you an empty database, not a working replacement, because there's no mechanism copying the actual data into it.
A failure mode redundancy alone doesn't protect against, for the stateful case: data loss or corruption on the primary itself. If the primary's disk corrupts a row, or a bad write silently corrupts application-level data, having a redundant (but not yet caught-up, or synchronously replicating that same bad write) standby doesn't help, because either the standby doesn't have the data yet (async lag) or it faithfully replicated the corruption along with everything else (synchronous replication of a logically bad write). Redundancy protects against a node dying; it does not protect against the data itself being wrong, that requires backups (a separate, point-in-time copy decoupled from live replication) and, for silent corruption specifically, checksums or application-level validation.
How this generalizes: a useful mental checklist for fault tolerance covers five distinct techniques, and redundancy and replication are only two of them: retries (recover from a transient failure by trying again), bulkheads (isolate one failure from spreading to unrelated resources), failover (the mechanism that switches traffic to a healthy replacement), redundancy (having that replacement exist at all), and replication (making sure the replacement actually has the state it needs). A strong answer names which of these a given design decision is actually addressing, since "redundancy" gets used loosely to mean all five in casual conversation.
Trade-offs & pitfalls
- Replication has a cost redundancy alone doesn't: network bandwidth, storage for extra copies, and a consistency model to reason about (synchronous replication costs write latency; asynchronous replication risks data loss on failover, the classic RPO trade-off).
- A common wrong turn: assuming "we have 3 replicas" automatically means "we're protected," without checking replication lag. A replica that's minutes behind at failover time silently loses however much data arrived in that window, unless the promotion logic explicitly accounts for lag and refuses to promote a too-far-behind replica.
- Redundancy for stateless services is comparatively cheap and low-risk to over-provision; replication for stateful services is not, since more replicas means more write-path coordination overhead (for synchronous replication) or more divergence risk (for asynchronous/multi-leader), so it isn't a "just add more" lever in the same way.
Give me an example of a time you received tough feedback or criticism right after something went wrong operationally, like after an outage. How did you manage your reaction in the moment, and what did you do afterward to rebuild trust?
Sample Answer
Direct answer
In the moment, my first job is to actually listen to the criticism rather than start explaining or defending myself before I've fully heard it, even when the instinct to justify is strong. Afterward, rebuilding trust isn't about the conversation where I received the feedback, it's about visibly acting differently going forward in the specific way the feedback pointed at.
Structured elaboration
- Managing the reaction in the moment: the instinct right after an outage, already stressed, is to explain the context and mitigating factors as soon as criticism starts. I've learned to let the person finish first, genuinely hear the specific complaint, and only then respond, since jumping in early to explain often lands as defensiveness even when that isn't the intent.
- Separating the valid signal from the delivery: tough feedback right after an outage often arrives with real frustration attached. The useful move is extracting the actual substance, what specifically should have gone differently, rather than reacting to the tone it arrived in.
- Not over-apologizing either: there's a version of managing the reaction that overcorrects into excessive self-criticism, which doesn't address the substance any better than defensiveness does; the goal is a level, accurate acknowledgment, not performing contrition.
- Rebuilding trust afterward: the actual trust repair happens in what changes afterward, doing the specific thing the feedback pointed at differently next time, not in how gracefully the original conversation went.
Worked example
Right after an outage I'd contributed to, my manager gave me direct, pointed feedback in a one-on-one: that I'd been slow to escalate once it became clear I was stuck, and that the delay had made the outage longer than it needed to be. My first instinct was to explain the reasoning that had made sense to me in the moment, that I'd thought I was close to a fix. I held off on that and let them finish first, and once I actually listened past my own defensiveness, the specific point was fair: I had, in fact, kept trying alone for longer than made sense given how the situation was unfolding.
I acknowledged the specific point directly rather than the vaguer "I hear you, I'll do better," and said what I'd concretely do differently: escalate earlier next time I'm stuck past a set point, rather than continuing to push alone. The actual trust rebuilding happened over the incidents that followed, not in that conversation. In the very next incident where I got stuck, I escalated well before I would have previously, and I made a point of telling my manager afterward that I'd deliberately applied the earlier feedback, which is what actually closed the loop for them, seeing the specific behavior change rather than just hearing that I'd taken the feedback well.
Trade-offs and pitfalls
The common failure mode is treating receiving feedback well as the whole task, being gracious and non-defensive in that one conversation and considering it handled. Without a visible change in behavior afterward, gracious listening reads as agreeable in the moment and forgotten a week later, which damages trust more than a defensive reaction followed by real change would. The other trap is swinging to excessive self-criticism, which can feel like taking it seriously but doesn't actually engage with the specific, actionable substance of the feedback any better than dismissing it does.
A user reports they cannot authenticate with their SSH public key. The auth log shows: 'sshd[1234]: Authentication refused: bad ownership or modes for directory /home/bob'. Explain the root causes for this message and provide the exact commands you would run to fix the problem and verify SSH key authentication will work. Also mention additional checks (e.g., SELinux) you would perform if permissions look correct.
Sample Answer
Direct answer
That specific message means the SSH daemon's (sshd's) strict permission checks rejected bob's home directory or .ssh directory before it ever evaluated the key itself. By default, sshd refuses to trust a public key if the account's home directory, its .ssh directory, or the key file is writable by anyone other than its owner, because a group- or world-writable path there would let another local user plant their own key and log in as bob. The fix is tightening ownership and permissions on the exact chain of directories sshd checks, then verifying the fix actually took effect.
Structured elaboration
Root causes. Any of these produce the same "bad ownership or modes" message:
/home/bobis not owned bybob(or root), often from files restored by a backup tool or copied in by another administrator's account./home/bobor/home/bob/.sshhas a group- or world-writable bit set (sshd'sStrictModessetting, on by default, rejects this)./home/bob/.ssh/authorized_keysitself is writable by group or other, even if the directories above it are fine.
Commands to fix it, applied in the order that matches how sshd walks the path:
# 1. Home directory must be owned by bob, and not writable by group or other.
sudo chown bob:bob /home/bob
sudo chmod go-w /home/bob
# 2. .ssh directory: owner-only access.
sudo mkdir -p /home/bob/.ssh
sudo chown bob:bob /home/bob/.ssh
sudo chmod 700 /home/bob/.ssh
# 3. authorized_keys file: owner-only read/write, no execute needed.
sudo chown bob:bob /home/bob/.ssh/authorized_keys
sudo chmod 600 /home/bob/.ssh/authorized_keys
Verifying the fix. Confirm the actual mode bits, not just that the commands ran without error:
stat -c "%a %U:%G %n" /home/bob /home/bob/.ssh /home/bob/.ssh/authorized_keys
Expect something like 755 bob:bob /home/bob, 700 bob:bob /home/bob/.ssh, and 600 bob:bob /home/bob/.ssh/authorized_keys. Then attempt the actual login with verbose logging so sshd's own reasoning is visible: ssh -vvv bob@host, and separately watch the server side with sudo tail -f /var/log/auth.log (or /var/log/secure on Red Hat-family distributions) for the corresponding "Accepted publickey" line, rather than assuming success from the client side alone.
If permissions look correct and it still fails, check SELinux (Security-Enhanced Linux, a mandatory access control system available on many Linux distributions). Standard Unix permission bits and SELinux are two independent checks; a directory can have perfectly correct ownership and mode and still be denied if its SELinux file context (label) is wrong, most commonly after files were restored from a backup or copied from another location without preserving context:
ls -Z /home/bob/.ssh /home/bob/.ssh/authorized_keys
.ssh and its contents should carry the ssh_home_t context. If it shows something else (commonly a generic context like user_home_t after a restore), restore the correct label rather than trying to set it by hand:
sudo restorecon -Rv /home/bob/.ssh
If that does not resolve it, check for an actual denial being logged rather than guessing further:
sudo ausearch -m avc -ts recent
An access vector cache (AVC) denial entry naming sshd and the .ssh path confirms SELinux, not a stale permission bit, is the actual blocker.
Worked example
A concrete before-and-after for this exact scenario. Suppose /home/bob/.ssh was left group-writable (mode 775) after a provisioning script ran as root and did not lock it down, a common way this exact error gets introduced:
Before: drwxrwxr-x bob bob /home/bob/.ssh (group-writable, sshd rejects this)
-rw-r--r-- bob bob /home/bob/.ssh/authorized_keys
After applying the fix commands above:
drwx------ bob bob /home/bob/.ssh
-rw------- bob bob /home/bob/.ssh/authorized_keys
Running the fix commands against exactly this before-state (verified directly: chmod go-w on the home directory, chmod 700 on .ssh, chmod 600 on authorized_keys) produces exactly the after-state shown, the group-write bit is gone from .ssh and authorized_keys no longer has any bits set beyond owner read/write. That is the state sshd's strict-mode check requires before it will even look at the key's contents.
Trade-offs and pitfalls
- Fixing only
authorized_keysand skipping the parent directories is a common half-fix. sshd checks the whole path, home directory,.ssh, and the key file itself, so a perfectly permissioned key file inside a group-writable.sshdirectory still fails. - Assuming a fix worked because the commands returned no error is a common mistake. Always verify the actual mode bits with
stat, and confirm the login itself succeeds with verbose logging, rather than trusting thatchmodsilently doing nothing wrong means the state is now correct. - Jumping straight to SELinux before confirming ownership and mode bits wastes time. SELinux denials produce a similar practical symptom (login refused) but a different log signature; check the simpler, more common cause first, and only chase an SELinux label mismatch once the standard permission chain is confirmed correct.
- Manually setting an SELinux context instead of using
restoreconrisks getting the label subtly wrong.restoreconrestores the context the system's own policy defines for that path, which is more reliable than guessing a context string by hand.
Explain Kubernetes' maxSurge and maxUnavailable rolling-update parameters. For a 100-pod deployment, how would you configure them differently for a latency-sensitive service versus a batch worker?
Sample Answer
Direct answer
maxSurge controls how many EXTRA pods above your target replica count can exist temporarily during a rollout; maxUnavailable controls how many pods can be DOWN (below target) at once. Together they set the trade-off between rollout speed/extra resource usage and how much capacity you're willing to sacrifice mid-rollout.
Structured elaboration
For a 100-pod deployment:
maxSurge: 25%means up to 25 extra pods can be created (125 total running temporarily) before old ones start being removed.maxUnavailable: 25%means up to 25 pods can be unavailable at once, so at minimum 75 pods are always serving traffic.- Setting both to 0 isn't allowed (Kubernetes needs at least one of them to make forward progress), and setting both to a large value maximizes rollout SPEED at the cost of both extra resource consumption AND reduced available capacity simultaneously.
Latency-sensitive service: prefer maxUnavailable: 0 (never drop below full target capacity, since capacity loss directly hurts p99 latency under load) combined with a modest maxSurge (say 10-25%) so new pods come up gradually without demanding a huge resource spike all at once. This is slower but protects the SLA.
Batch worker: can tolerate maxUnavailable being higher (say 50%), since a temporarily smaller worker pool just means jobs queue a bit longer rather than users seeing degraded latency; this lets the rollout complete faster with less need for surge capacity, since you're willing to sacrifice throughput temporarily instead of paying for extra headroom.
Worked example
For the latency-sensitive service at 100 replicas with maxUnavailable: 0, maxSurge: 25: Kubernetes first creates 25 new pods (125 total), waits for them to become ready, then removes 25 old pods (back to 100, all new), repeats: creates 25 more new (125 again, this time mixed), removes 25 more old, and so on, always keeping at least 100 healthy pods serving traffic throughout. For the batch worker at maxUnavailable: 50, maxSurge: 0: Kubernetes removes up to 50 old pods first (down to 50 capacity), creates 50 new ones to replace them, and repeats, needing no extra capacity headroom at all but running at half capacity partway through.
Trade-offs and pitfalls
Higher surge means faster rollouts but a temporary resource spike your cluster's autoscaler or node capacity needs to be able to absorb; if it can't, new pods stall in Pending and the rollout stalls too. The common mistake is picking these values once and never revisiting them as traffic patterns or cluster capacity change, so a setting that was safe at launch quietly becomes risky (or needlessly slow) as the service scales.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Systems Administrator jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs