FAANG-Standard Interview Preparation Guide: Junior Systems Administrator
This guide is based on general FAANG interview practices and may not reflect specific company procedures.
FAANG-standard interview process for junior systems administrators typically consists of 7 rounds spanning 4-6 weeks. The process emphasizes technical depth in operating systems and infrastructure fundamentals, practical troubleshooting ability, system administration hands-on experience, and cultural alignment. Junior-level candidates are expected to demonstrate solid foundational knowledge, independent competency on routine infrastructure tasks with occasional guidance, and the ability to learn quickly in complex technical environments. The interview process progressively assesses technical breadth, hands-on capabilities, behavioral traits, and role fit.
Interview Rounds
Recruiter Screening Call
What to Expect
Initial conversation with a technical recruiter to assess background, motivation, and basic technical foundation. The recruiter will verify your resume details, understand your interest in systems administration, confirm your willingness to work on-call or during maintenance windows if required, and assess your communication skills. They may ask high-level technical questions to verify you're at the appropriate level. This round is typically conversational and designed to ensure mutual fit before investing in technical interviews.
Tips & Advice
Be enthusiastic about infrastructure and systems work. Have a clear story about why you chose systems administration as a career. Be honest about your experience level - recruiters appreciate candidates who know their strengths and gaps. Prepare 2-3 concrete examples of infrastructure challenges you've solved or situations where you learned something valuable. Ask thoughtful questions about the team structure, on-call rotation, and growth opportunities. Research the company's public infrastructure approach or engineering blog beforehand to show genuine interest.
Focus Topics
Role Expectations and Fit
Understanding of systems administrator responsibilities, on-call requirements, shift patterns, and typical challenges. Realistic expectations about junior-level work (routine tasks, guided projects, learning opportunities).
Practice Interview
Study Questions
Communication and Professionalism
Ability to articulate technical concepts clearly, answer questions concisely, and demonstrate professional demeanor. Shows whether you can communicate effectively with non-technical stakeholders and teammates.
Practice Interview
Study Questions
Technical Foundation Assessment
High-level questions about operating systems (Linux/Windows), networking concepts, and infrastructure basics. Recruiter may ask 'What is DNS?' or 'Explain the difference between TCP and UDP?' to gauge baseline knowledge.
Practice Interview
Study Questions
Career Motivation and Background
Clear articulation of why you chose systems administration, your relevant experience (academic projects, personal labs, internships, entry-level roles), and specific technologies you've worked with. Ability to describe your technical journey and what excites you about infrastructure work.
Practice Interview
Study Questions
Technical Screen - Linux and Operating Systems Fundamentals
What to Expect
First technical interview conducted by a senior systems administrator or infrastructure engineer. Focus is on Linux/Unix fundamentals, operating system concepts, command-line proficiency, and basic system administration tasks. This is typically a conversational technical interview with whiteboarding or screen-sharing where you'll be asked to explain concepts, work through scenarios, and demonstrate Linux command knowledge. You may be asked to design simple solutions or troubleshoot basic Linux problems.
Tips & Advice
Get very comfortable with Linux command-line - this is non-negotiable for systems administrators. Practice common commands daily (ls, grep, find, sed, awk, systemctl, etc.). Understand the filesystem hierarchy and permissions model deeply. When asked a question, think out loud and explain your reasoning - interviewers want to see your thought process, not just the answer. If you don't know something, admit it honestly and discuss how you would find the answer. Bring examples of systems you've managed or labs you've built. Practice explaining OS concepts (processes, memory, file systems) clearly. Be prepared for 'walk me through' questions like 'Walk me through booting a Linux system' or 'Walk me through how permissions work'.
Focus Topics
Package Management and Software Updates
Linux package managers: apt/apt-get for Debian-based systems, yum/dnf for Red Hat-based systems. Understanding package dependencies, version management, repository configuration, and security updates. Practical experience with installing, updating, and removing packages.
Practice Interview
Study Questions
Operating System Concepts and Process Management
Understanding processes, threads, process lifecycle (fork, exec, exit, wait), process states (runnable, sleeping, zombie), process scheduling, CPU and memory management. Familiarity with ps, top, htop, nice, renice commands. Understanding load average, context switching, and resource utilization.
Practice Interview
Study Questions
Networking Fundamentals for Systems Administrators
Basic networking concepts: OSI model layers, TCP/IP stack, DNS, DHCP, routing basics. Network configuration on Linux: ifconfig/ip command, network interfaces, static vs. dynamic IP configuration, routing tables, netstat/ss for network diagnostics. Understanding common ports and services (HTTP:80, HTTPS:443, SSH:22, DNS:53).
Practice Interview
Study Questions
Linux Command-Line and Shell Scripting
Proficiency with essential Linux commands (grep, find, sed, awk, cat, less, head, tail, ps, top, systemctl, journalctl). Understanding of shell scripting basics including variables, conditionals, loops, and function writing. Familiarity with bash, environment variables, and command substitution. Ability to write simple bash scripts for common administrative tasks.
Practice Interview
Study Questions
Linux File System, Permissions, and User Management
Understanding of Linux file system hierarchy (/etc, /var, /usr, /home, /opt, etc.), file permissions (rwx for user/group/other), special permissions (setuid, setgid, sticky bit), and umask. User and group management including adding/removing users, managing groups, sudo privileges, and understanding /etc/passwd, /etc/shadow, and /etc/sudoers.
Practice Interview
Study Questions
Technical Screen - Windows Server and Active Directory Administration
What to Expect
Second technical interview focusing on Windows Server administration, Active Directory (AD) concepts, user account management, and Group Policy. This interview is conducted by a Windows infrastructure specialist or systems engineer. You may be asked scenario-based questions like 'A user can't log in - how would you troubleshoot?' or 'How would you implement a password policy across 500 computers?' Expect discussion of Windows Server architecture, domain concepts, and practical administrative tasks.
Tips & Advice
If you have primarily Linux background, invest significant time studying Active Directory and Windows Server basics. Understand the difference between workgroups and domains. Study Group Policy deeply - it's a major mechanism for Windows administration at enterprise scale. Know how to navigate Active Directory Users and Computers MMC console conceptually. Practice explaining Windows concepts clearly. Be honest about Windows experience if it's limited - you can still demonstrate understanding of concepts. Prepare examples of Active Directory issues you've helped troubleshoot or labs you've built. Understand Windows user account types and permissions model. Study common Windows Server roles: DNS, DHCP, File Server, Print Server basics.
Focus Topics
Windows Server Roles and Services
Common Windows Server roles: Active Directory Domain Services (AD DS), DNS Server, DHCP Server, File and Storage Services, Print Server, Remote Desktop Services basics. Understanding server manager and how to add/remove roles. Purpose and function of key services.
Practice Interview
Study Questions
Group Policy Fundamentals
Understanding Group Policy Objects (GPOs), Group Policy Editor, and how policies apply to users and computers. Policy hierarchy and inheritance (OU-level policies override parent policies). Security filtering and WMI filters. Common policy settings: password policies, audit policies, software restriction policies. Troubleshooting GPO application issues using gpresult and event logs.
Practice Interview
Study Questions
Windows System Administration Tools and Diagnostics
Essential Windows administration tools: Computer Management MMC, Active Directory Users and Computers, Group Policy Editor (gpedit.msc), Services (services.msc), Event Viewer. Command-line tools: ipconfig, nslookup, net commands, Get-ADUser (PowerShell), dsquery. Understanding event logs and performance monitoring.
Practice Interview
Study Questions
Windows User Account and Permission Management
Creating and managing user accounts, resetting passwords, enabling/disabling accounts. Group management: creating security groups, assigning users to groups, nested groups. NTFS permissions: read, write, modify, full control. Permission inheritance and effective permissions. Local administrators group and domain admin groups. UAC (User Account Control) basics.
Practice Interview
Study Questions
Active Directory Architecture and Concepts
Understanding Active Directory structure: forests, trees, domains, organizational units (OUs). Domain controllers, global catalog, LDAP basics. User objects, computer objects, groups (security groups vs. distribution groups). Domain membership and authentication flow. Trust relationships between domains. Ability to navigate AD Users and Computers and understand AD hierarchy.
Practice Interview
Study Questions
Technical Assessment - Infrastructure Design and System Services
What to Expect
This round evaluates your ability to think about infrastructure design, backup strategies, system monitoring, and disaster recovery at a basic level. This is not a distributed systems design interview, but rather a practical infrastructure assessment. You may be asked: 'Design a backup strategy for a small company with 50 servers', 'How would you set up monitoring for a data center?', or 'Walk me through your approach to patching a critical vulnerability across 100 systems.' Conducted by an infrastructure or operations engineer, this assesses your understanding of operational concerns beyond individual machine administration.
Tips & Advice
Remember this is junior-level - don't over-engineer solutions. Focus on practical, straightforward approaches with clear reasoning. Discuss trade-offs (cost vs. redundancy, automation vs. manual process) and justify your choices. Ask clarifying questions about requirements (budget, scale, SLA requirements). Draw diagrams if helpful. Discuss monitoring and alerting - these are critical for operations. Mention backup verification and disaster recovery testing. For junior level, the interviewer wants to see understanding of concepts and reasonable problem-solving, not expert-level infrastructure design. Walk through your thinking process clearly. If you're unsure about a concept, discuss what you would research or who you would consult. Provide examples from systems you've actually managed.
Focus Topics
Infrastructure Scalability and Capacity Planning
Understanding capacity planning basics: current utilization, growth trends, forecasting. Scaling approaches: vertical scaling (bigger hardware) vs. horizontal scaling (more systems). Load balancing concepts. Storage growth management. Planning for future growth without over-provisioning. Cost optimization in infrastructure planning.
Practice Interview
Study Questions
Patch Management and Software Updates
Patch management strategy: prioritizing critical vs. non-critical patches, testing procedures, deployment windows. Understanding patch impact and rollback procedures. Automation tools for patch deployment. Balancing security updates against operational stability. Patch verification and compliance tracking.
Practice Interview
Study Questions
System Monitoring and Alerting
Monitoring architecture: agents vs. agentless monitoring, metric collection, log aggregation. Key metrics: CPU, memory, disk utilization, network bandwidth. Alerting: threshold-based alerts, alert fatigue, escalation procedures. Log analysis and centralized logging. Common monitoring tools and their purposes. Performance baselines and anomaly detection.
Practice Interview
Study Questions
Infrastructure Security and Hardening
Security hardening basics: disabling unnecessary services, applying security patches promptly, strong authentication (passwords, multi-factor authentication), principle of least privilege. Firewall configuration concepts. Security audit logs. Compliance requirements (HIPAA, PCI-DSS basics). Incident response procedures and security incident escalation.
Practice Interview
Study Questions
Backup and Disaster Recovery Strategy
Understanding backup strategies: full backups, incremental/differential backups, backup frequency decisions. Recovery time objective (RTO) and recovery point objective (RPO). Off-site backup storage and 3-2-1 backup rule (3 copies, 2 different media, 1 off-site). Backup verification and restore testing. Disaster recovery plan basics: failover procedures, business continuity planning. Common backup solutions and tools.
Practice Interview
Study Questions
Practical Hands-On Assessment
What to Expect
Live technical assessment where you solve practical infrastructure problems in a lab environment or through simulated scenarios. You may be given access to virtual machines (Linux and/or Windows) where you perform tasks such as: configuring a user account with specific permissions, troubleshooting a connectivity issue, setting up basic monitoring, configuring a backup, or troubleshooting a service that won't start. Alternatively, this may be scenario-based where you're asked step-by-step how you would approach specific problems. This round directly assesses your hands-on capability and practical problem-solving approach.
Tips & Advice
Stay calm and methodical - speed is less important than approach at junior level. Ask clarifying questions about the objective and constraints. For troubleshooting scenarios, work through problems systematically: understand the symptom, check logs, review configuration, test hypotheses. Document your steps as you go. Explain what you're doing and why - this helps interviewers understand your thought process. If you get stuck, discuss what you would try next or who you would consult. Practical labs test your ability to execute, not memorization. Practice hands-on labs beforehand using free resources like VirtualBox and Linux VMs. Be comfortable using command-line and GUI tools. If given a task you're unsure about, ask for clarification rather than guessing. Show willingness to try different approaches if one doesn't work.
Focus Topics
Documentation and Communication During Tasks
Clearly explaining what you're doing and why as you work through problems. Taking notes and documenting steps. Communicating findings and recommendations clearly. Asking clarifying questions when unclear. Demonstrating understanding of what you're doing, not just executing commands.
Practice Interview
Study Questions
Systematic Troubleshooting Methodology
Approaching problems methodically: gathering information about the symptom, checking error messages and logs, developing hypotheses about root cause, testing hypotheses systematically, documenting findings and solutions. Knowing when to escalate issues vs. resolve independently. Requesting help appropriately without giving up too quickly.
Practice Interview
Study Questions
Hands-On Linux System Administration
Practical tasks on Linux systems: user account creation and management, file permission configuration, package installation and updates, service management (start/stop/restart services), log file analysis, basic troubleshooting of common issues (connectivity, service failures, disk space). Navigation of Linux filesystem and system configuration files.
Practice Interview
Study Questions
Hands-On Windows Server Administration
Practical Windows tasks: creating and managing user accounts in Active Directory, setting NTFS permissions on folders, applying Group Policy settings, managing services, configuring network settings, troubleshooting common Windows issues. Navigation of Windows administrative tools and command-line utilities.
Practice Interview
Study Questions
Behavioral and Soft Skills Assessment
What to Expect
Interview focused on behavioral competencies, soft skills, and cultural fit. Interviewer will ask about teamwork, communication, learning ability, handling pressure, and conflict resolution using STAR method questions (Situation-Task-Action-Result). Expect questions like: 'Tell me about a time you made a mistake and how you handled it', 'Describe a situation where you had to learn a new technology quickly', 'Tell me about a time you had to communicate a technical issue to a non-technical person'. This round assesses whether you fit the team culture and have the interpersonal skills for successful collaboration.
Tips & Advice
Prepare 5-7 specific examples using the STAR method covering different situations (challenges overcome, mistakes learned from, collaboration, learning agility, prioritization under pressure). Be authentic and honest - interviewers can tell when you're not being genuine. Focus on what YOU did, not what your team did. For junior level, emphasize learning ability, willingness to help colleagues, and positive attitude. Discuss how you handle on-call rotation or after-hours emergencies positively. Share examples that demonstrate problem-solving approach, not just technical outcomes. Practice telling stories concisely - aim for 2-3 minutes per example. Prepare questions about team dynamics, how junior admins are mentored, growth opportunities, and technical challenges the team faces. Research the company's values and mission - reference them if authentic to your experience.
Focus Topics
Communication and Stakeholder Management
Explaining technical concepts to non-technical stakeholders (users, business managers). Documenting systems clearly for other team members. Listening skills and understanding requirements. Escalating issues appropriately. Status updates and progress communication.
Practice Interview
Study Questions
Attention to Detail and Quality Mindset
Examples of catching errors before they impact systems. Verification and testing procedures you follow. Double-checking critical changes. Documentation accuracy and completeness. Pride in work quality even for routine tasks.
Practice Interview
Study Questions
Collaboration and Teamwork
Examples of working effectively with teammates, asking for help when appropriate, helping junior colleagues or peers learn. Handling different communication styles and perspectives. Contributing to team knowledge base or documentation. Respecting expertise of more senior colleagues while contributing your own ideas.
Practice Interview
Study Questions
Handling Mistakes and Pressure
Examples of mistakes you made early in your career and how you recovered. Approach to preventing similar mistakes. Handling on-call or after-hours emergencies calmly. Prioritization when dealing with multiple urgent issues. Asking for help when situations exceed your capability.
Practice Interview
Study Questions
Learning Agility and Adaptability
Demonstrated ability to learn new technologies and systems quickly. Examples of situations where you learned something new outside your comfort zone and applied it successfully. Growth mindset and willingness to tackle unfamiliar problems. How you approach learning new platforms or tools. Examples from academic projects, personal labs, or early career experiences.
Practice Interview
Study Questions
Hiring Manager Round
What to Expect
Final interview with the hiring manager or team lead who will directly supervise you. This round focuses on role fit, team dynamics, career expectations, and whether you understand the day-to-day realities of the position. The hiring manager will discuss the team structure, current infrastructure challenges, growth trajectory, mentorship expectations, and your career goals. They'll assess whether you're someone they can manage effectively and whether you'll thrive on their team. This is also your opportunity to ask detailed questions about role expectations, team culture, and career development.
Tips & Advice
This is a bidirectional conversation - the hiring manager wants to understand you AND you should thoroughly evaluate whether this role is right for you. Be honest about your experience level and learning goals. Ask thoughtful questions about team structure, mentorship, how junior admins are developed, technical challenges the team faces, on-call rotation expectations, and growth trajectory. Discuss your interest in specific technologies or areas of infrastructure. Be authentic about what excites you and what concerns you about the role. Show that you've done research on the company's infrastructure challenges or technical direction (if public information is available). Discuss how you work with seniors and your approach to learning on the job. Prepare questions about technical growth, career path options, team dynamics, and how successes are measured for junior admins.
Focus Topics
Career Growth and Development Path
Understanding progression from junior to mid-level admin role. Typical timeline for advancement. Required skills and experiences for promotion. Opportunities for specialization (security, database administration, cloud infrastructure, etc.). Continuing education and certification support.
Practice Interview
Study Questions
Technical Challenges and Interesting Projects
Current infrastructure challenges the team is working on. Recent infrastructure improvements or migrations. Technology stack and tools the team uses. Opportunities for junior admins to contribute to meaningful projects. Interesting technical problems the team is solving.
Practice Interview
Study Questions
Team Dynamics and Work Environment
Understanding team size, structure, and dynamics. How team members interact and support each other. On-call support procedures and how junior admins are integrated into on-call rotations. Team priorities and values (automation, documentation, security, etc.). Culture of continuous improvement versus maintaining status quo.
Practice Interview
Study Questions
Role Expectations and Day-to-Day Responsibilities
Clear understanding of typical daily responsibilities, on-call rotations, shift patterns, and how junior admin role differs from mid-level positions. Knowledge of current team projects and infrastructure challenges. Realistic expectations about routine versus complex tasks at junior level. Understanding of how work is prioritized and time allocated.
Practice Interview
Study Questions
Mentorship and Learning Opportunities
Understanding how junior admins are mentored and trained on the team. Availability of experienced senior admins for questions and guidance. Formal or informal knowledge transfer. Opportunities to work on increasingly complex projects. Access to training and certification opportunities. How the team approaches knowledge sharing.
Practice Interview
Study Questions
Frequently Asked Systems Administrator Interview Questions
Your company wants users' profiles and data to follow them across laptops, desktops and remote session hosts. Compare roaming profiles, folder redirection and container-based profile solutions, then recommend an approach for branch users on slow links and another for pooled virtual desktops, covering the trade-offs and storage you would plan for.
Sample Answer
Direct answer
Use Folder Redirection with Offline Files (Windows keeps a cached local copy of network files, so users can keep working when the link is down and changes sync back later) for branch users on slow links, keeping any roaming profile small and limited to their own computers, because a roaming profile is copied across the link at every sign-in and sign-out. Use FSLogix profile containers (a virtual hard disk file per user (VHD or VHDX) on a file share, attached at sign-in) for pooled virtual desktops, because there the profile must follow the user to whichever machine in the pool they land on and must not be copied at all. Size container storage from measured usage, not from the per-user ceiling.
The three approaches compared
| Roaming user profiles | Folder Redirection (with Offline Files) | Container-based profiles (FSLogix, user profile disks) | |
|---|---|---|---|
| How it works | The profile loads from a file share to the local computer at sign-in, merges with any local copy, and the local copy merges back to the server at sign-out | A known folder such as Documents points at a path on a file share; applications use it as if local | A virtual disk (VHD or VHDX) holding the profile is mounted at sign-in, and a filter driver (software that sits between applications and the file system and redirects file access) makes applications see a local profile |
| Network cost | Whole profile copied at sign-in and sign-out | Files fetched as used; Offline Files caches them for offline use | No copy; reads and writes go to the disk file |
| Sign-in time on a slow link | Grows with profile size | Small, since only settings roam | Small, as nothing is copied |
| Weak points | Large profiles, profile versions differ between Windows versions | Needs a reachable server, or Offline Files to hide an outage | Depends on file share availability and storage design |
| Best for | Small settings-only profiles on a few trusted computers | Document-centric users, branch offices, mobile users | Pooled or multi-session virtual desktops |
User profile disks (the Remote Desktop Services feature that keeps each user's profile in a VHDX file on a share) are also a container approach, and Microsoft's roaming-profile article points to them for Start menu roaming on session hosts. FSLogix is the current option for pooled desktops and is not limited to virtual desktops.
Branch users on slow links
Recommendation: Folder Redirection plus Offline Files for the data, and a roaming profile only if settings must follow users between computers. Both Folder Redirection and roaming profiles can be limited to a user's primary computers, using the Group Policy settings "Download roaming profiles on primary computers only" and "Redirect folders on primary computers only" (they depend on the msDS-PrimaryComputer attribute in Active Directory, an attribute on the user's account that lists the computers designated as that user's own). That keeps profile data off shared or conference-room machines.
The cost of getting this wrong is arithmetic. Assume a dedicated 10 Mbit/s branch link, a full-rate transfer and no protocol overhead:
10 Mbit/s500 MB×8=400 s10 Mbit/s50 MB×8=40 sA 500 MB roaming profile costs about 6.7 minutes at sign-in; redirecting the large folders so the profile is 50 MB brings that to about 40 seconds. The 500 MB and 50 MB figures are illustrative; measure your own profiles. Real transfers are slower, and sign-out repeats the copy in the other direction.
Storage and share design (from Microsoft's deployment guidance):
- Deploy Folder Redirection first on existing local profiles, then roaming profiles, so profiles start small.
- Keep the roaming profile share separate from shares used for redirected folders, to prevent inadvertent offline caching of the profile folder. The share's caching mode (
New-SmbShare -CachingMode) controls offline caching. - On a clustered file share, disable continuous availability (an SMB share setting that lets open files survive a failover to another cluster node) for roaming profiles to avoid performance problems. If the share sits behind Distributed File System (DFS) Namespaces, give the folder a single target, and if DFS Replication copies it, let users reach only the source server, or users will make conflicting edits.
- Do not place Folder Redirection, home directories or roaming profiles on a Scale-Out File Server; Microsoft marks them "not recommended" there because they generate many writes that must be committed immediately. Use a File Server for general use.
- If you support two Windows versions, keep separate profile versions per OS; each gets its own folder, so profiles double in count and storage, and changes made on one version do not roam to the other. Redirect common folders so files are visible on both.
- Use the permissions Microsoft lists for the profile share: System Full control; Administrators Full control on the folder only; Creator Owner (a built-in identity that stands for whoever created the item) Full control on subfolders and files only; the user group List folder and Create folders on the folder only; everything else removed. Access-based enumeration and Encrypt data access can be enabled on the share.
Pooled virtual desktops
Recommendation: FSLogix Profile Container for each user on Server Message Block (SMB) file shares, with everything in the single profile container. FSLogix also has a separate Office container (ODFC, which holds only Office data such as Outlook and OneDrive cache); it is optional and is mainly paired with another roaming-profile product.
Settings that matter (from the FSLogix configuration reference and examples):
| Setting | Value to use | Why |
|---|---|---|
Enabled | 1 | Required |
VHDLocations | One UNC (Universal Naming Convention, \\server\share) path to the share | Where containers live |
VolumeType | VHDX | Preferred over VHD for supported size and fewer corruption cases |
SizeInMBs | Default 30000 | The maximum container size; containers are extended to it at sign-in; it can be raised later but not lowered |
IsDynamic | Default 1 | Disk usage grows only as needed, up to SizeInMBs |
ProfileType | 0 (default) | A user's container is mounted through only one connection at a time, which keeps things simple and fast |
DeleteLocalProfileWhenVHDShouldApply | 1 | Prevents a stray local profile masking the container; Microsoft warns FSLogix permanently deletes the local profile in that case, so confirm nothing needed lives only in local profiles |
ClearCacheOnLogoff (Cloud Cache only) | 1 | Saves local disk on pooled machines |
Two of these settings are marked required in Microsoft's reference, so containers do not work without them: Enabled and VHDLocations (with Cloud Cache, CCDLocations takes the place of VHDLocations). VolumeType and SizeInMBs have defaults (vhd and 30000) that this design overrides or confirms, and the other rows tune behaviour or apply only when Cloud Cache is used.
Availability: listing several paths in VHDLocations does not give resilience. Microsoft describes the list as locations to search for the user's container, and if none holds one a new container is created in the first listed location, so a user whose existing container is not found in any searched location gets a new, empty profile. Microsoft's high-availability article also states that standard containers provide no resiliency of their own and rely entirely on the storage provider, and each container still lives in exactly one location, so a lost share still stops the users whose containers are on it. For resilience use Cloud Cache (an FSLogix option that keeps a local cache and writes each container to two or more storage providers, meaning separate storage locations such as file shares or Azure storage accounts; CCDLocations lists them, with at least two), and set HealthyProvidersRequiredForRegister to 1 as in Microsoft's example so a user cannot start a local cache when a provider is down. Do not change FlipFlopProfileDirectoryName, SIDDirNameMatch, SIDDirNamePattern, VHDNameMatch or VHDNamePattern on an existing deployment, because those five settings control how FSLogix names and finds each user's folder and disk, so after a change it looks for the old containers under new names, will not find them and will create new empty profiles.
Currency note: Microsoft's FSLogix pages warn that a Windows change in the April 2026 Windows Server update, still worded as upcoming on those pages although that month has now passed, moves the default Kerberos encryption type (Kerberos is the Windows sign-in ticket protocol) from RC4, an older cipher, to AES-SHA1, a stronger encryption and integrity combination, and that file shares hosting containers that have not been upgraded to AES-SHA1 might have access problems afterwards. The link to profiles: the session host reaches the container share over SMB and proves who the user is with Kerberos tickets, so a change in the default ticket encryption can break access to the share unless the storage also supports AES-SHA1. Confirm the storage accounts and file servers hosting containers support AES-SHA1 before that update is installed, and if the session hosts already have it, check now.
Storage planning for 500 users at the default ceiling:
500×30000 MB=15000000 MBConverting that needs one stated convention. Treating a megabyte as 1,048,576 bytes (so the unit is really a mebibyte, MiB), one container is 30000 / 1024 = 29.3 GiB and the total is 15,000,000 / 1,048,576 = 14.3 TiB. Treating a megabyte as 1,000,000 bytes, the total is 15,000,000 x 1,000,000 bytes = 15 TB. These are two readings of the same 15,000,000 "MB" figure, not one quantity in two units: 15,000,000 x 1,048,576 bytes is 15.7 TB (14.3 TiB), while 15,000,000 x 1,000,000 bytes is 15 TB (13.6 TiB), a gap of about 5 percent, so state which convention you use when you order storage. Either way it is the worst case when every container reaches its maximum, not the expected usage. Because dynamic containers grow only as used, run a pilot, measure the actual size of each container, and plan on a high percentile plus headroom. For illustration, if the measured average were 8 GiB, the demand would be 500 x 8 = 4,000 GiB, which is 4,000 / 1,024 = 3.9 TiB, well below the 14.3 TiB ceiling. Use the ceiling for alerts and quotas, and use measured data for purchases. Also plan antivirus exclusions (listed in the FSLogix prerequisites) and test the share and NTFS (New Technology File System) permissions before the pilot.
Pitfalls
- Using roaming profiles for pooled desktops: sign-in and sign-out time grow with the profile, because the profile is copied each time.
- Putting containers on a share that is not highly available and calling it done; a share outage then stops every sign-in.
- Letting the working data live in the profile. Redirect or keep documents in a separate store so profile size stays bounded.
- Assuming a container's size limit can be shrunk later; it cannot, so start with a deliberate ceiling.
Explain bufferbloat: why excessive buffering in network devices increases latency and jitter under load even though it reduces packet loss, and how Active Queue Management algorithms such as fq_codel counteract it. Why does bufferbloat specifically interfere with TCP's own congestion signals?
Sample Answer
Direct answer
Bufferbloat happens when routers or switches along a path have excessively large buffers that queue packets during congestion instead of dropping them, which keeps loss low but lets queuing delay grow essentially unbounded, adding latency and jitter that TCP's own congestion control never gets a clear enough signal to react to. Active Queue Management algorithms like fq_codel counteract it by proactively dropping or marking packets BEFORE the queue grows large, restoring a timely loss signal.
Structured elaboration
A classic loss-based congestion-control algorithm relies on packet loss (or an explicit congestion mark) as its primary "back off" signal. If a device's buffer is very large, it can absorb a burst of excess traffic by queuing it rather than dropping it, so loss never actually happens, but every packet sitting in that oversized queue now experiences extra delay waiting its turn. Ironically, a device built to be MORE forgiving (bigger buffer, less loss) ends up making the user experience WORSE (much higher and more variable latency) precisely because it hides the congestion signal the sender needs to slow down.
fq_codel (Fair Queuing with Controlled Delay) attacks this from two angles: "fair queuing" gives each active flow its own small queue so one bulk-transfer flow can't monopolize the buffer and starve a latency-sensitive flow (like a video call) sharing the same link; "controlled delay" tracks how long packets are actually sitting in the queue, and once a packet has been queued longer than a target delay (commonly around 5ms), it starts dropping packets to force the responsible flow's congestion control to back off, well before the queue grows large enough to cause serious latency.
Worked example
On a home internet connection with a large, un-managed buffer at the router, a single large upload (like a cloud backup) can push queuing delay from a normal few milliseconds up to several hundred milliseconds or more, which is directly noticeable as a video call over the same connection becoming choppy or laggy, even though no packets for the video call are actually being dropped, they're just sitting in a queue behind the bulk upload's traffic for hundreds of milliseconds. Enabling fq_codel on that router's egress queue caps how long any packet can wait, keeping the video call's latency low even while the bulk upload continues in the background.
Trade-offs & pitfalls
Bufferbloat is easy to misdiagnose as "the link doesn't have enough bandwidth" when the real problem is excess, unmanaged queuing delay on a link that has PLENTY of bandwidth; the fix (AQM, not more bandwidth) is the opposite of what a naive read of "things feel slow" would suggest. A useful diagnostic: latency under load (with a saturating background transfer running) that's dramatically higher than idle-latency is the signature of bufferbloat, not a raw throughput problem.
Describe the core components and responsibilities of a Public Key Infrastructure (PKI): root and intermediate CAs, certificate issuance, certificate validation (CRL and OCSP), certificate lifecycle management, and why short-lived certificates reduce risk. Include operational challenges of running an internal CA at scale.
Sample Answer
Direct answer
A Public Key Infrastructure, or PKI, is the set of roles, policies, and systems that let you issue, distribute, and verify digital certificates, which bind a public key to an identity (a person, service, or device) so others can trust that a given public key really belongs to that name. At its center sits a hierarchy of Certificate Authorities, or CAs: a root CA whose trust is simply assumed (its certificate is pre-installed as trusted), and one or more intermediate CAs it delegates to, which do the day-to-day issuing so the root's highly sensitive private key can stay offline and rarely used.
Structured elaboration
Root and intermediate CAs. The root CA's certificate is self-signed, and its public key is the actual trust anchor: everything else is trusted only because it chains back to the root. Because the root's private key is the single most valuable secret in the whole system (its compromise makes every certificate ever issued under it suspect), best practice keeps it offline, powered on only rarely to sign a new intermediate certificate, and delegates all routine issuance to intermediate CAs, whose own certificates are signed by the root. This creates a chain of trust: leaf certificate, then intermediate CA certificate, then root CA certificate, and a verifier walks this chain, checking each link's signature, before trusting the leaf.
flowchart TD
Root[Root CA certificate: self-signed, kept offline] --> Intermediate[Intermediate CA certificate]
Intermediate --> Leaf[Leaf certificate: e.g. a server or service identity]
Certificate issuance. An entity generates a key pair and submits a CSR (certificate signing request) containing its public key and identity claims, such as a domain name, to the CA. The CA validates the claimed identity (for a public web CA, this is domain-control validation; for an internal CA, it might be validating that the request came from an authorized internal system) and, if satisfied, signs a certificate binding that public key to that identity, with a defined validity period.
Certificate validation: CRL and OCSP. A certificate can become untrustworthy before its stated expiry, for example if the private key leaks, so verifiers need a way to check whether it's been revoked.
- CRL (certificate revocation list). The CA periodically publishes a signed list of every revoked certificate's serial number, and a verifier downloads and checks it. Simple, but the list can get large, and it's only as fresh as its publication interval, leaving a window where a just-revoked certificate can still check out fine.
- OCSP (Online Certificate Status Protocol). A verifier asks the CA, or a designated OCSP responder, in real time whether a specific certificate is still valid, and gets back a signed yes, no, or unknown answer. This is fresher than a CRL, but adds a network round trip and a dependency on the responder's availability. OCSP stapling, where the server itself periodically fetches its own OCSP response and presents it alongside its certificate, removes the client's need to call out at all, fixing both the latency and the privacy leak of the responder learning who's checking which certificate.
Certificate lifecycle management. The lifecycle runs: issuance, then distribution and installation, then monitoring for approaching expiry, then renewal (a new certificate issued before the old one expires), then revocation if the certificate is compromised or no longer needed, then expiry. At scale this needs automation, since tracking thousands of certificates' expiry dates by hand doesn't work; this is exactly why protocols like ACME (Automatic Certificate Management Environment), which let a server request and renew its own certificate programmatically, exist.
Why short-lived certificates reduce risk. A certificate's validity window is exactly the window during which a leaked private key stays directly useful to an attacker without anyone having to actively revoke anything. A one-year certificate leaked on day two gives an attacker roughly a year of usable window if revocation checking isn't happening or fails silently; a 24-hour certificate leaked at the same point gives at most hours. Shortening the lifetime shifts the safety property from "we must catch and revoke every compromise" (which depends on detection, and on every verifier actually checking revocation, which, as shown above, has real gaps with both CRL and OCSP) to "compromises age out on their own," a much stronger guarantee that depends far less on humans noticing in time. The cost is that short-lived certificates require fully automated issuance and renewal, since a human manually renewing a certificate every 24 hours isn't viable, which is exactly why short lifetimes and automation tend to appear together.
Operational challenges of running an internal CA at scale.
- Trust distribution. Every machine or service that needs to trust your internal CA's certificates must have that CA's root certificate installed in its own trust store; rolling this out, and later rotating it, across a large, heterogeneous fleet of different operating systems, containers, and legacy systems is itself a distribution problem.
- Availability. If the CA, or its issuing path, goes down, nothing new can get a certificate issued or renewed. For a short-lived-certificate architecture this becomes a hard dependency: an outage longer than your certificate lifetime cascades into fleet-wide expiry failures, meaning the CA's own availability requirement is often higher than that of the services depending on it.
- Revocation infrastructure. Standing up and keeping available a CRL or OCSP responder that scales with your fleet size and query volume.
- Key protection for the CA's own private keys. The same HSM-backed (hardware security module), access-controlled, audited practices used for the root, scaled to however many intermediates you run.
- Monitoring and alerting for approaching expiries, as a safety net even with full automation, since automation itself can silently fail, for example a renewal job quietly stopping for one service with nobody noticing until certificates start expiring.
Worked example
A browser connects to a server presenting a leaf certificate for internal-svc.example.com, signed by an internal intermediate CA, Corp Issuing CA 2026, which is itself signed by the offline Corp Root CA. Verification walks the chain: the leaf's signature is checked against the intermediate's public key (valid), then the intermediate's own certificate signature is checked against the root's public key (valid, and the root is present in the trust store), so the chain is trusted. Separately, the verifier checks whether the leaf has been revoked, for example via OCSP: it asks the OCSP responder about the leaf's serial number and gets back a signed "good" response, valid for the next hour, a typical OCSP response validity window. Had the response instead said "revoked," the chain-of-trust check above would be irrelevant; the certificate is rejected regardless of how clean its signature chain looks.
Trade-offs and pitfalls
Skipping revocation checking entirely, which many client libraries default to for performance or availability reasons, silently reintroduces exactly the risk that short-lived certificates and revocation both exist to solve, and it's a common blind spot.
Treating the root CA like any other key, keeping it online and hot, defeats the whole point of the hierarchy; a root compromise is catastrophic precisely because everything else chains back to it.
Assuming "internal" means "less rigor than a public CA": an internal CA that's easier to compromise than the systems it protects becomes a single high-value target, which is a worse security posture than not having internal PKI at all.
Underestimating CA availability requirements once short-lived certificates are adopted: an outage that would be a minor inconvenience with year-long certificates becomes a production emergency with 24-hour certificates.
Your organisation is adopting Microsoft 365 and needs sign-in with on-premises AD while keeping password policies. Compare the ways to connect the two directories in terms of sign-in behaviour, security, operational overhead and resilience.
Sample Answer
Direct answer
Default to password hash synchronization (PHS) plus Seamless single sign-on (SSO), which signs in domain-joined PCs without a password prompt. It has the least infrastructure, keeps working when on-premises servers are down, and still applies your on-premises password complexity and history rules at change time. Choose pass-through authentication (PTA, where an on-premises agent checks the password against a domain controller) only if you must enforce on-premises lockout, disabled, expired and sign-in-hours state at every single sign-in. Choose federation with Active Directory Federation Services (AD FS) only for needs Microsoft Entra ID cannot meet natively, such as third-party multifactor authentication (MFA) or sign-in with a sAMAccountName. Enable PHS even if you choose PTA or federation, as the fallback.
What each method does at sign-in
A few terms first. A tenant is your organization's own dedicated instance of Microsoft Entra ID. Microsoft Entra Connect is the on-premises software that copies users from AD to the tenant. Federation means the tenant hands the password check to a separate on-premises sign-in service (AD FS), which vouches for the user afterwards.
- Password hash sync (PHS): Connect copies a processed form of each password hash to the cloud, and the cloud itself checks the password. Nothing on-premises is involved at sign-in.
- Pass-through authentication (PTA): the cloud passes the typed password to an on-premises agent, which asks a domain controller whether it is right and returns yes or no.
- Federation (AD FS): the cloud redirects the user to the on-premises AD FS farm, which checks the password and issues a signed token back to the cloud.
For choosing a method, what matters is where the password is checked, what happens when on-premises is down, and how quickly a disabled account is refused; the intervals and agent counts in the table are operating details that follow from that choice.
The options compared
| Password hash sync (PHS) | Pass-through authentication (PTA) | Federation (AD FS) | |
|---|---|---|---|
| Where the password is checked | In the cloud, against a hash of a hash synchronized from AD | On-premises, by an agent that validates against a DC | On-premises, by the federation farm |
| On-premises policy behaviour | Complexity and history apply when the password changes on-premises; expiry in the cloud is "never" by default for synced users; disabled accounts can lag up to 30 minutes; locked-out and expired states are not synced | Disabled, locked out, account expired, password expired and sign-in hours are enforced at each sign-in | Same state checks as PTA, enforced on-premises |
| Seamless SSO | Yes | Yes | No, it cannot be used with AD FS |
| Security | Entra holds a salted hash of the AD password hash (derived with PBKDF2, a deliberately slow hashing function, so it is a "hash of a hash"), which cannot be replayed in a pass-the-hash attack on-premises (an attack that signs in with a stolen hash instead of the password) | Password validation never leaves the premises; agents need unconstrained access to DCs, so they cannot sit in a perimeter network | Largest on-premises attack surface (web application proxy servers in the perimeter, certificates) |
| Operational overhead | Lowest: part of the sync process, runs every 2 minutes and the interval is not configurable | Agents on existing servers; three recommended | Highest: two or more AD FS servers and two or more proxies, TLS certificates, health monitoring |
| Resilience | Cloud service; survives an on-premises outage | Needs agents and DCs reachable; failover to PHS is manual through Entra Connect | Farm and DCs must be up; PHS can be a fallback |
Behaviour detail
- Sign-in with PHS. A user changing a password on-premises signs in with it after the next sync, normally minutes. Existing cloud sessions are unaffected. Bulk account disables should be followed by an immediate sync cycle.
- Seamless SSO creates an
AZUREADSSOcomputer account in each forest. Microsoft recommends rolling its Kerberos decryption key (the secret that lets Entra decrypt the Kerberos tickets browsers present) at least every 30 days withUpdate-AzureADSSOForest, once per forest only, because Microsoft documents that running it more than once per forest stops the feature until users' Kerberos tickets expire and are reissued by AD. It works only with PHS or PTA. - Password expiry. If synced users must follow cloud password-expiry rules, enable the
CloudPasswordPolicyForPasswordSyncedUsersEnabledfeature before enabling PHS; accounts that must never expire (service accounts) then needDisablePasswordExpirationset explicitly.
High availability of the sync engine
Microsoft Entra Connect Sync can only be active-passive. A second server in staging mode (a standby copy kept current but not allowed to write to the cloud) imports and synchronizes but does not export, and does not run password sync or password writeback until you disable staging. If it has been in staging for long, password sync needs a catch-up period after it goes active, and newly changed passwords do not work in the cloud until the backlog clears. Staged rollout, mentioned in the pitfalls, is a different feature: it moves chosen user groups to cloud authentication gradually. Check state on either server:
Import-Module ADSync
Get-ADSyncScheduler | Select-Object StagingModeEnabled
Microsoft Entra Connect cloud sync is the lighter alternative, because it does not depend on a single sync server and can provide higher availability. For PTA, install three agents (the first with Connect plus two more) so one can be down for maintenance while another fails.
Writeback
Password writeback (which keeps the on-premises password current after a cloud change) sends a password changed or reset in the cloud (self-service password reset or an administrator in the Entra admin center) back to AD DS. It works with PHS, PTA and federation, enforces your on-premises history, complexity, age and filters, uses only outbound port 443, and cannot reset passwords for protected-group members. Only one active Connect server can use writeback at a time, and it is not reliable with staged rollout enabled for a security group.
Worked example
At 09:00 HR disables a leaver's AD account. Under PTA the leaver's next sign-in is refused, because the agent checks the live account state. Under PHS the disabled state can lag up to 30 minutes, so the leaver may still sign in until about 09:30 unless you force a sync cycle after the change. If that window is unacceptable for a high-risk population, that is the case for PTA, with PHS still enabled as the fallback.
Trade-offs and pitfalls
- Switching methods needs planning; staged rollout lets you move users gradually.
- PTA keeps password validation on-premises but adds an on-premises dependency; PHS is the opposite.
- Federation is the most capable and the most fragile.
Your org has a major initiative with dependencies across product, design, data, and engineering, but each function has different priorities and limited capacity. Walk me through how you would align the groups, identify trade-offs, and create a plan everyone can commit to.
Sample Answer
I’d start by aligning everyone on the outcome, not the function-specific asks.
Step 1: Clarify the shared goal
I’d bring product, design, data, and engineering into one working session and define the business outcome, success metrics, and deadline constraints.
Step 2: Map dependencies and capacity
I’d list the critical dependencies, identify who owns each one, and make capacity visible by function. That exposes where the real bottlenecks are.
Step 3: Sequence the plan
I’d build the plan around the critical path: what must happen first, what can run in parallel, and what can be deferred. If capacity is tight, I’d use a simple trade-off framework: highest business value, lowest risk, and strongest dependency unlocks first.
Step 4: Create commitment
I’d confirm decision rights, document what each team is committing to, and define checkpoints where we can re-plan if assumptions change.
The goal is not to make everyone equally happy; it’s to make the trade-offs explicit so each group can commit to a plan they helped shape.
Worked example
Say the initiative is a checkout redesign that needs a payments-data migration (data team), a new UI (product design and frontend), and an updated fraud-detection model (data science). In the working session, the shared goal turns out to be reducing checkout abandonment by a set amount before the next major sales event, which becomes the deadline constraint. Mapping dependencies shows the new UI can't ship until the data migration completes, and the fraud model needs at least two weeks of production traffic on the new UI before it can be retrained safely, so the data migration is the critical-path item. Applying the trade-off framework, the data migration (highest dependency-unlock value) is sequenced first, the UI ships second, and the fraud-model update is explicitly deferred to just after the sales event rather than rushed; each team commits to that sequence in writing, with a checkpoint two weeks before launch to re-plan if the migration slips.
Users can pass a filename to your script. Write the part that accepts it and guarantees it cannot escape a base directory, including when symlinks are involved.
Sample Answer
Direct answer
Turn the user's name into a canonical absolute path first (every .. collapsed and every symlink replaced by what it points to), then check that the canonical result sits under the canonical base directory. A check on the raw text, such as "reject names containing ..", is not enough, because a symlink that already lives inside the base can point anywhere. Because the file can change between the check and the open (a time-of-check to time-of-use race, TOCTOU), open the file and then check again what the open descriptor really refers to.
The artifact
This was run as root inside a throwaway ubuntu:24.04 container (GNU coreutils), because it creates /srv directories. resolve_in_base does the check; open_and_recheck opens the checked path and re-checks the open file; read_in_base runs the two in order. Two pieces of Bash syntax appear in the code: set -u makes the script stop with an error if it uses a variable that was never set (so a typo cannot silently become an empty string), and $'\n' is Bash quoting that produces a real newline character.
#!/usr/bin/env bash
set -u
# Prints the canonical path of $2 inside base directory $1, or fails.
resolve_in_base() {
local base=$1 name=$2 root target
[[ -n $name ]] || { echo "empty name" >&2; return 1; }
[[ $name != /* ]] || { echo "absolute path refused: $name" >&2; return 1; }
[[ $name != *$'\n'* ]] || { echo "newline in name refused" >&2; return 1; }
root=$(realpath -e -- "$base") || return 1
[[ $root != / ]] || { echo "base must not be /" >&2; return 1; }
target=$(realpath -m -- "$root/$name") || return 1
[[ $target == "$root"/* ]] || { echo "escapes base: $name -> $target" >&2; return 1; }
printf '%s\n' "$target"
}
# Opens an already-checked path, then re-checks what the descriptor really points at.
open_and_recheck() {
local base=$1 target=$2 root real status
root=$(realpath -e -- "$base")
exec 3< "$target" || return 1
real=$(readlink -f /proc/self/fd/3)
if [[ $real != "$root"/* ]]; then
exec 3<&-
echo "descriptor escaped base: $real" >&2
return 1
fi
cat <&3
status=$? # keep the read's own status; closing below would overwrite it
exec 3<&-
return "$status"
}
read_in_base() {
local target
target=$(resolve_in_base "$1" "$2") || return 1
open_and_recheck "$1" "$target"
}
base=/srv/uploads
mkdir -p "$base/reports" /srv/uploads-evil /srv/secret
echo "quarterly numbers" > "$base/reports/q1.txt"
echo "evil sibling" > /srv/uploads-evil/x.txt
echo "root password" > /srv/secret/shadow
echo "attacker copy of q1" > /srv/secret/q1.txt
ln -s /srv/secret/shadow "$base/reports/link-to-file"
ln -s /srv/secret "$base/reports/link-to-dir"
ln -s /srv/secret/not-yet "$base/reports/dangling-out"
ln -s q1.txt "$base/reports/alias-inside"
for n in "reports/q1.txt" "reports/alias-inside" "reports/new file.txt" \
"../uploads-evil/x.txt" "reports/../../secret/shadow" "/srv/secret/shadow" \
"reports/link-to-file" "reports/link-to-dir/shadow" "reports/dangling-out" \
"-rf" ""; do
printf '%-30s => ' "[$n]"
resolve_in_base "$base" "$n" 2>&1
done
echo "--- read"
read_in_base "$base" "reports/q1.txt"
read_in_base "$base" "reports/link-to-file"
read_in_base "$base" "reports"; echo "directory read status: $?"
echo "--- swap between the check and the open"
target=$(resolve_in_base "$base" "reports/q1.txt")
echo "checked: $target"
mv "$base/reports" "$base/reports.old" # attacker swaps the directory ...
ln -s /srv/secret "$base/reports" # ... for a symlink to elsewhere
open_and_recheck "$base" "$target"
Output:
[reports/q1.txt] => /srv/uploads/reports/q1.txt
[reports/alias-inside] => /srv/uploads/reports/q1.txt
[reports/new file.txt] => /srv/uploads/reports/new file.txt
[../uploads-evil/x.txt] => escapes base: ../uploads-evil/x.txt -> /srv/uploads-evil/x.txt
[reports/../../secret/shadow] => escapes base: reports/../../secret/shadow -> /srv/secret/shadow
[/srv/secret/shadow] => absolute path refused: /srv/secret/shadow
[reports/link-to-file] => escapes base: reports/link-to-file -> /srv/secret/shadow
[reports/link-to-dir/shadow] => escapes base: reports/link-to-dir/shadow -> /srv/secret/shadow
[reports/dangling-out] => escapes base: reports/dangling-out -> /srv/secret/not-yet
[-rf] => /srv/uploads/-rf
[] => empty name
--- read
quarterly numbers
escapes base: reports/link-to-file -> /srv/secret/shadow
cat: -: Is a directory
directory read status: 1
--- swap between the check and the open
checked: /srv/uploads/reports/q1.txt
descriptor escaped base: /srv/secret/q1.txt
Reading the re-check, line by line
A file descriptor is a small number the kernel gives a process for each file it has open; the process uses the number as a handle instead of the name. Descriptors 0, 1 and 2 are already taken (standard input, output and error), so 3 is the first free one, and nothing more special than that.
exec 3< "$target"opens the file for reading and attaches it to descriptor 3 for the rest of the script. If the open fails,|| return 1stops./proc/self/fd/3is Linux's window onto the process's own open files:/procis a virtual directory the kernel fills in,selfmeans the process doing the looking, andfd/3is a symlink to whatever descriptor 3 is open on. Here the looking is done byreadlink, a child process started by$(...); it inherits descriptor 3 from the script (descriptors opened withexecare inherited by child programs), so its/proc/self/fd/3is the very same open file.readlink -ffollows it and prints the real canonical path, sorealis the file that was actually opened, not the name you asked for.[[ $real != "$root"/* ]]is the same trailing-slash comparison as before, now applied toreal. If it fails,exec 3<&-closes descriptor 3 (<&-means "close"), the function reportsdescriptor escaped base, and returns 1 before any data is read.cat <&3reads from descriptor 3 (the already-checked file) and prints it. Its exit status is saved instatusbefore the finalexec 3<&-closes the descriptor andreturn "$status"hands it back; without that, theexecwould overwrite the status with 0 and a failed read would look like success. The demo's last read shows it: reading the directoryreportsmakescatprintcat: -: Is a directory, and the function returns 1 rather than 0.
Why each line is there
realpath -e -- "$base"canonicalizes the base. GNUrealpath -erequires every component to exist, so a mistyped base fails instead of silently matching nothing.realpath -m -- "$root/$name"canonicalizes the target.-mallows missing components, so a file you are about to create (reports/new file.txt) can be validated, while symlinks that do exist are still resolved. That is whylink-to-file,link-to-dir/shadowand even the danglingdangling-outlink (its target does not exist yet, but it would be created outside the base) are all refused, whilealias-inside, a link that stays inside the base, is allowed and resolves toq1.txt.[[ $target == "$root"/* ]]has a trailing slash on purpose./srv/uploads-evil/x.txtstarts with the characters/srv/uploads, but not with/srv/uploads/, and the sibling-directory row shows it being refused. The quoted"$root"makes any glob characters in the base literal; the unquoted/*is the pattern.--afterrealpath, and joining the name onto the base, mean a name like-rfis a file called-rf, never an option.- Absolute names are refused as a policy: joined onto the base they would just be read as relative, so refusing makes an attack attempt visible instead of quietly "working". Names containing a newline are refused because
$(...)strips trailing newlines, which would make the checked path differ from the path the user meant. - Empty names and a base of
/are refused, since/would make every path "inside".
The remaining race, and what actually closes it
Between resolve_in_base returning and the open, anyone who can write inside the base could swap a directory for a symlink. read_in_base closes that for reads: after exec 3< "$target" it asks the kernel where descriptor 3 really points (readlink -f /proc/self/fd/3, Linux only) and refuses unless that is under the base. The check is on the file that was opened, so a swap after the check cannot be used.
The last block of the demo plays that attack out with real paths. resolve_in_base approves /srv/uploads/reports/q1.txt (a real file inside the base). Before the open, the attacker renames the reports directory and puts a symlink named reports pointing at /srv/secret in its place. The path string is unchanged, so the open succeeds, but it now lands on /srv/secret/q1.txt. The first check passed on the old layout; the descriptor check asks the kernel what descriptor 3 really points at, gets /srv/secret/q1.txt, and refuses with descriptor escaped base: /srv/secret/q1.txt. This runs as root in a throwaway container, so the attacker's mv and ln are simulated by the script itself.
That does not help a write or create, because by the time you can inspect the result the data has landed. For writes, remove the attacker's ability to plant symlinks (the process that fills the base writes regular files only, and untrusted users cannot create links there), or, outside Bash, do the open in a language that can call openat2 with RESOLVE_BENEATH, which makes the kernel refuse any path resolution that leaves the directory (Linux 5.6 and later; RESOLVE_NO_SYMLINKS refuses symlinks entirely). Bash has no way to pass those flags, so this is a note about other languages, and the realpath plus descriptor re-check above is the approach for a Bash script.
Trade-offs and pitfalls
realpath -mandreadlink -fare GNU/Linux behaviour; macOS and BusyBox differ, so say which platform the script targets.- Opening a path that an attacker has swapped for a FIFO (a named pipe) blocks until a writer appears: in a test, a read of a FIFO inside the base hung until a 3-second
timeoutkilled it. The descriptor check cannot help because the open itself never returns, so wrap reads of a shared base intimeoutor refuse non-regular files before opening. - A symlink loop does not make
realpath -mfail, it just returns a path; the later open of such a path fails, which is the safe outcome. - Do not "sanitize" by stripping
../withsed:....//becomes../after one pass, and it ignores symlinks completely. - A canonical-path check follows symlinks but cannot see hard links. A hard link is a second name for the same file data, so a hard link inside the base to a sensitive file looks like any other file in the base and is allowed. Prevent that by controlling who can create links in the base.
You are on-call and receive alerts that a production web application is returning 500 errors and experiencing high latency from multiple regions. Describe, step‑by‑step, a systematic troubleshooting process you would follow to identify the root cause. Include: what data and artifacts you would collect first, which commands/tools you would run on affected hosts, how you'd triage service vs network vs DB vs infrastructure, and how you'd prioritize actions under time pressure.
Sample Answer
A strong candidate starts from a fixed loop, not a guess: gather signal, form a hypothesis, run the cheapest test that could disprove it, and only then act.
The loop
- Collect first, don't touch anything. Pull the alert/report, recent deploys and config changes, and the three signal types: metrics (what changed and when), logs (what the service says happened), traces (where time went across components).
- Triage the layer before the cause, using a fixed set of checks per layer (see below): does it correlate with one host, one region, one dependency, or all of them? A failure isolated to one instance points at that instance; a failure across all instances after a deploy points at the deploy; a failure correlated with a dependency's own error rate points downstream.
- Form one falsifiable hypothesis at a time ("the new deploy is the cause") and pick a test that would disprove it cheaply (check whether the errors started at the exact deploy timestamp, or whether rolling back one canary host clears them). Avoid changing five things at once.
- Prioritize under time pressure: mitigate first (rollback, scale out, fail over) if user impact is ongoing, investigate the true cause in parallel or after.
Commands and tools per layer, and how to triage between them
Run these roughly in parallel across a couple of affected hosts, not sequentially one host at a time:
- Service/application layer:
kubectl get pods -o wideandkubectl describe pod <pod>(orsystemctl status <service>on VMs) to check restart counts and recent events;kubectl logs -f <pod> --previousorjournalctl -u <service> -ffor the exact error at the moment of the alert;top/htoporps aux --sort=-%cpu,-%memfor CPU/memory pressure on the process itself; for managed runtimes,jstack <pid>orjcmd <pid> Thread.printto check for stuck threads. Signal that this is the layer: errors and restarts correlate with specific pods/hosts or with the deploy timestamp, not with a single dependency or the network path. - Network layer:
curl -vagainst the exact endpoint from an affected host to separate DNS/TLS/connect time from application response time;dig/nslookupfor DNS resolution issues;traceroute/mtrfor path/latency between hops;ss -sornetstat -antpfor socket/connection-state saturation (too manyTIME_WAIT, exhausted ephemeral ports);tcpdumpon a specific host if a particular hop is suspected. Signal:curl -v's connect/TLS phase is slow while the app's own processing time (visible in traces) is normal, or the problem tracks a specific region/CDN edge rather than a specific service version. - Database layer: active/slow query lists (
SHOW PROCESSLISTon MySQL,pg_stat_activityon Postgres), lock waits (SHOW ENGINE INNODB STATUS,pg_locks), and the app's own connection-pool metrics (checked-out connections near the pool limit). Signal: request latency traces show most of the time inside the DB span, and DB-side query/lock metrics show a corresponding spike at the same timestamp. - Infrastructure layer:
kubectl describe node/kubectl top nodefor node-level CPU/memory/disk pressure,dmesgorjournalctl -kfor OOM-killer or kernel-level events, and the cloud provider's status page or recent autoscaling/capacity events. Signal: the failure correlates with a specific node, availability zone, or a capacity/autoscaling event rather than with a code deploy or a single dependency.
The layer whose checks show a signal exactly aligned with the alert's onset time is the one to dig into first; the others should still be glanced at briefly to rule out a compounding factor, but don't get equal depth until the primary layer is ruled out.
Worked example
A service starts returning 500s at 14:02. A deploy went out at 14:00. Metrics show error rate flat on hosts still running the old build and elevated only on hosts running the new one. That single comparison (same traffic, different build, different outcome) is strong evidence for the deploy as root cause, tested in under a minute using data you already have, before touching any code.
Trade-offs and pitfalls
The most common mistake is skipping straight to "it's probably the database" because that's where the last incident was, without checking whether this failure actually correlates with DB latency. A hypothesis not tested against data is a guess wearing an RCA costume. The other common failure is fixing the first plausible thing that appears in the logs, when it's a symptom of an earlier upstream cause; correlating the alert time against the deploy/change timeline first avoids that trap.
What is the difference between a playbook, a runbook and a README? For each, say who reads it, what it contains, how often it needs maintenance, and a situation where you would choose it over the other two.
Sample Answer
Direct answer. All three are documents, but they answer different questions. A README answers "what is this and how do I start?", a runbook answers "what exact steps do I follow for this task or alert?", and a playbook answers "how do we handle this kind of situation, and who decides what?". Naming conventions differ between companies, so I state my definitions when I write one. In everyday terms: a runbook is an operating instruction sheet, like a recipe you follow exactly; a playbook is a team's game plan, which tells people their roles and when to make a call; a README is the front-door sign of a project.
| README | Runbook | Playbook | |
|---|---|---|---|
| Reader | New developer or user opening the repository | The on-call engineer or operator under time pressure | A team or incident commander (the person coordinating a response) plus other roles |
| Contents | What the project is, install and run, how to contribute, where to find more docs | Step-by-step commands for one procedure (restart a service, fail over a database), with checks and rollback | Higher-level plan for a scenario type: roles, decision points, escalation, communication, links to runbooks |
| Maintenance | Whenever setup or usage changes (on each pull request that alters them); reviewed each release | Every time the procedure or system changes, plus a periodic test (for example, quarterly, by running it) | After every real use or exercise, and at least yearly |
| Choose it when | You want a stranger to run the project in ten minutes | An alert fires and the response is a known, repeatable procedure | The situation needs judgement and coordination, such as a security breach or a regional outage |
Worked example: a payments service.
- README: "Payments API. Requires Python 3.12. Run
make dev. Tests:make test. Docs: link." - Runbook: "Alert: queue depth above threshold. 1. Check the consumer dashboard. 2. Run
kubectl rollout restart deployment/payments-consumer. 3. Confirm the queue drains within 10 minutes. 4. If not, roll back to the previous release (command)." Each step is unambiguous and checkable. - Playbook: "Suspected payment data leak. Incident commander declares severity. Security lead assesses scope. Legal decides on notification. Communications drafts customer message. Decision point: take the service offline? criteria listed." It links to the runbooks for isolating a host.
Choosing between them. If the reader must think, it is a playbook; if the reader must follow, it is a runbook; if the reader is arriving cold, it is a README. A typical mistake is writing a runbook full of judgement calls (it fails at 3am) or a playbook full of shell commands (it goes stale quickly).
Production service across a fleet shows increased latency and monitoring indicates high iowait on Linux hosts. Provide a step-by-step troubleshooting plan using Linux tools to determine if the cause is device saturation, kernel/block-layer contention, application IO patterns, or cloud storage latency. Include commands and short-term mitigations.
Sample Answer
Direct answer
The four candidate causes leave different fingerprints: cloud storage latency hits every host sharing that
backing volume or tier at the same time with normal per-process CPU; device saturation or kernel/block-layer
contention shows up as elevated %util/await on one or a few local devices independent of which
application is running; and an application I/O pattern problem shows up as one specific process dominating
pidstat -d on an otherwise-healthy device. Work through that ordering rather than guessing, and apply the
matching short-term mitigation once you've localized it.
Step-by-step plan
- Confirm scope: is the elevated
iowait(the percentage of CPU time cores spent idle while I/O was
outstanding) truly fleet-wide, or a subset of hosts sharing one backing store, availability zone, or
storage tier? A dashboard slicing p95iowaitby host and by storage backend answers this in one look
and immediately points toward "shared infrastructure" versus "one bad host or one bad app". - Rule out device saturation and kernel/block-layer contention on each affected host:
iostat -x 1for
%utilandawaitper device;vmstat 1'sst(steal time, CPU cycles a hypervisor took away from
this VM) column, since high steal can produce I/O-shaped symptoms that are actually a noisy-neighbor CPU
problem, not storage at all;dmesg/journalctl -kfor multipath failover events orblk_update_request
retries indicating the block layer itself is unhappy, not just busy. - Is it disk, filesystem, or a specific process: run
pidstat -d 1on an affected host. If one process
dominates the I/O and the device isn't otherwise saturated, that's an application I/O pattern issue
(fix the app or its schedule). If the device itself shows high%util/awaitroughly proportional to
every process's normal I/O, the bottleneck is the storage layer, not any one application. Cross-check
withlsofon the busy mount or device to see exactly which files or sockets are held open by whatever
pidstatflagged. - Cloud storage latency specifically: compare the
awaitvalues against the provisioned performance
tier's expected latency for that volume type, and check for burst-credit exhaustion (many cloud block
storage tiers grant a burst allowance above a lower sustained baseline; once burst credits run out,
awaitrises sharply with no change in application behavior at all). If several hosts on the same
backing store all show the sameawaitjump at the same time with unrelated workloads, that is strong
evidence for storage-layer latency rather than any one host's application.
Worked example
A PostgreSQL read replica in a fleet of database hosts starts showing 40% iowait. pidstat -d 1 shows the
postgres WAL (write-ahead log) writer process with an elevated write rate, but pg_stat_activity and the
application's own metrics confirm the actual write VOLUME hasn't changed. iostat -x 1 on the data volume
shows await jumped from a 2ms baseline to 45ms while throughput (wkB/s) stayed essentially flat, the
signature of latency degrading with no corresponding increase in demand, which points at the storage layer
rather than the application. Checking the cloud provider's per-volume burst-credit metric confirms the
balance had hit zero. Short-term mitigation: redirect read traffic away from that replica to a healthier one
while requesting a volume-type upgrade (a higher baseline IOPS tier) rather than trying to tune the
application, since the application's behavior never changed.
Trade-offs and pitfalls
Don't conflate "iowait is high" with "the disk is the bottleneck": iowait is an accounting artifact of
idle CPU cores waiting for outstanding I/O, so on a busy multi-core host it can look deceptively low even
when one process is genuinely I/O-bound, always cross-check with per-device await, not iowait percentage
alone. Short-term mitigations differ by cause and picking the wrong one wastes the incident window: throttle
or reschedule the offending job for an application I/O pattern problem, fail over or scale reads away for a
storage-layer latency problem, but never run fsck, a remount, or any other write-adjacent "fix" as a
mitigation, that risks making an already-degraded host worse.
When you are dropped into a system you do not know, how do you decide whether to work it out on your own or go and ask someone? Walk me through how you make that call and what pushes it one way or the other.
Sample Answer
Direct answer
The call comes down to three things: how urgent the situation is, how much damage a wrong guess could cause, and how much of the answer is actually discoverable on my own versus locked in someone's head. When the blast radius is small and the information is findable, I work it out myself; when either the stakes are high or the knowledge simply isn't written down anywhere I can reach, I ask, and I try to ask well rather than asking instead of trying.
What pushes the decision each way
Toward figuring it out alone: low stakes if I'm wrong, a reversible action, and real evidence I can search, like existing code, logs, or documentation, even if imperfect. I'd rather spend twenty minutes tracing something myself than interrupt someone for a question the system can actually answer.
Toward asking: anything with real blast radius if I get it wrong, anything time-sensitive where figuring it out alone would blow a deadline that asking wouldn't, and anything that lives only in a person's head with no written trace, since no amount of my own digging will surface knowledge that was never recorded anywhere.
I also weigh whose time is actually being spent either way. Struggling alone for an hour on something a five-minute answer would resolve isn't more virtuous, it's just a worse use of everyone's time, mine included, once you account for the risk of getting it wrong.
A short illustration each way
I once spent about thirty minutes tracing through a configuration file to understand a setting rather than asking, because getting it wrong would have been low-stakes and immediately obvious if wrong, and I learned something about the system I'd have missed by just being told the answer. A different time, on a system with production traffic, I hit a setting I didn't understand within the first hour on a team, and I asked immediately rather than experimenting, because a wrong guess there could have affected real users, and there was someone two seats away who could tell me in thirty seconds what would have taken me an unknown amount of digging to maybe find.
Trade-offs and pitfalls
The pitfall on one end is interrupting people constantly for things you could find yourself, which costs their time and slows down your own ability to build real familiarity with the system. The pitfall on the other end is treating asking as a failure and pushing through alone on something high-stakes, which is how avoidable mistakes happen in systems you don't yet understand well enough to know what you don't know.
Recommended Additional Resources
- Linux Academy or A Cloud Guru Linux fundamentals courses
- Microsoft Learn: Windows Server Administration fundamentals
- Active Directory documentation and official Microsoft guides
- RHEL 8/9 and Ubuntu server administration documentation
- CompTIA A+ certification study materials (covers PC hardware and troubleshooting fundamentals)
- CompTIA Server+ certification materials (specifically covers server administration)
- 'The Practice of System and Network Administration' by Limoncelli (classic systems administration guide)
- 'Windows Server 2022 Administration' books and official Microsoft documentation
- Linux man pages and command-line documentation (man 5 sudoers, man 5 passwd, etc.)
- Cybrary and Udemy courses on Linux and Windows Server administration
- SANS Institute resources on systems administration (free tier available)
- Practical hands-on labs: VirtualBox with Linux VMs and Windows Server evaluation editions
- YouTube technical deep-dives on systems administration from trusted channels
- Red Hat Academy resources for Red Hat systems
- Canonical (Ubuntu) official training materials
- Your target company's engineering blogs and published infrastructure insights
- Stack Exchange and ServerFault for troubleshooting real-world scenarios
- Official vendor documentation: RedHat, Canonical, Microsoft, VMware (for virtualization)
Search Results
100+ AWS Interview Questions and Answers (2026) - Simplilearn.com
There are a number of different AWS-related questions covered in this article, ranging from basic to advanced, and scenario-based questions as well.
Operating System Interview Questions - GeeksforGeeks
Operating System Interview Questions · 1. What is a process and process table? · 2. What are the different states of the process? · 3. What is a Thread? · 4. What ...
42 HR Administrator Interview Questions and Sample Answers
HR administrator interview questions include: "Why are you interested in this role?", "What are your core strengths?", "What do you know about this position?", ...
50 Most Popular Salesforce Interview Questions & Answers ...
1. Describe how Salesforce CRM is used by organizations? At its core, Salesforce is a customer-facing CRM system. It is used to record customer ...
STAR Method Interview Questions & Answers - Interviews Chat
Explore top STAR Method interview questions and answers across a variety of roles, designed to help you ace your next interview with confidence.
What Is a Network Administrator? A Career Guide - Coursera
Interview questions for network administrator jobs · What is a firewall, and how would you implement one? · What is a proxy server? · What is a switch? · What types ...
This interview preparation guide was generated using AI-powered research from the sources listed above. While we strive for accuracy, we recommend verifying critical information from official company sources.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Systems Administrator jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs