Technical Discovery & Needs Qualification Questions
The diagnostic front half of a technical sale: how a presales or solutions professional uncovers a prospect's business objectives, pain points and technical environment, and qualifies the opportunity. Covers preparing for and structuring discovery calls and workshops, building credibility and handling guarded or skeptical technical contacts, discovery questioning and frameworks (SPIN, MEDDIC), mapping stakeholders, decision criteria, champions and the approval path, capturing non-functional requirements (performance, security, compliance, availability, integration), mapping the customer's current-state architecture, integrations and network or operational constraints, verifying budget, timeline and measurable success criteria, running a quick feasibility check on custom asks, handling infeasible or conflicting requirements found in discovery, and packaging findings into a discovery deliverable or CRM record. Stops short of building the ROI case or negotiating terms; eliciting product requirements from end users and internal stakeholders outside a sale is covered elsewhere.
A customer's requirements document says only that 'the system must be highly available'. Turn that into several measurable requirements you would propose to the customer, say what each would mean in contractual language, and explain how each ties to business impact.
Sample Answer
Direct answer
"Highly available" is not testable, so I would convert it into a small set of measurable targets, each with a number, a measurement method, a contractual term and a business-impact reason. I would propose them as a package: availability percentage, recovery time objective, recovery point objective, performance under load, and maintenance and support commitments. I would also ask the customer to rank them, because every extra nine costs money. A "nine" is one more 9 in the percentage: going from 99.9% to 99.99% is one extra nine and allows ten times less downtime.
Terms: availability is the share of time the service works as agreed. SLA (service-level agreement) is the contractual promise, usually with remedies. SLO (service-level objective) is the internal target behind it. RTO (recovery time objective) is how long an outage may last. RPO (recovery point objective) is how much recent data may be lost. P95 (95th percentile) latency is the response time that 95% of requests beat.
The measurable requirements
| Requirement | Example target | Contractual wording (sketch) | Business impact tie |
|---|---|---|---|
| Availability | 99.9% monthly for the production service | "Service Availability measured per calendar month as the percentage of minutes the service is up, excluding agreed maintenance windows; service credits (a refund or bill discount the vendor owes when the target is missed) apply below 99.9%." Maintenance windows are pre-announced periods when the service may be down without counting against the target. | Cost of one hour of downtime during trading |
| Recovery time (RTO) | 4 hours for a regional failure (an outage of a whole cloud region, meaning one geographic cluster of data centres) | "Service will be restored within 4 hours of a declared disaster." | Time the business can run on manual process |
| Recovery point (RPO) | 15 minutes of data | "No more than 15 minutes of committed data lost in a disaster." | Number of orders or records that would have to be re-entered |
| Performance under load | P95 response under 500 ms at 200 requests per second | "The 95th-percentile response time of requests to the production service shall not exceed 500 ms in any 5-minute window in which the request rate does not exceed 200 per second, measured at the service boundary." | Abandonment, SLA to the customer's own customers |
| Maintenance and notice | Planned maintenance at most 4 hours per month, with 5 days notice, outside 06:00-22:00 local | Scheduled-maintenance clause | Operational planning, avoiding peak periods |
| Support response | Severity 1 acknowledged in 30 minutes, 24x7 | Support schedule | How fast a human starts working the outage |
What the availability numbers mean (computed in a 30-day month)
Downtime allowed per 30-day month is (100 minus availability) percent of 43,200 minutes:
- 99.9% allows 43.2 minutes.
- 99.95% allows 21.6 minutes.
- 99.99% allows 4.32 minutes.
Per year (365 days, 525,600 minutes), 99.9% allows 525.6 minutes, about 8.8 hours. Series dependencies (components that must all work for the service to work) multiply: two components each at 99.9% in a chain give 0.999 x 0.999 = 99.8001%, about 99.8%, which allows roughly 86 minutes a month. I use that to show the customer that a headline number for the service is lower than for each part.
Business impact conversation
I ask: "What does one hour of downtime cost you in revenue or manual work, and what is the worst time of day or year?" If an hour costs 50,000 dollars at peak, then the 43.2 minutes allowed at 99.9% is about 50,000 x 43.2 / 60 = 36,000 dollars of exposure per month. At 99.95% the allowance is 21.6 minutes (18,000 dollars) and at 99.99% it is 4.32 minutes (3,600 dollars). Whether the extra nines are worth their price is then a comparison the customer can make, and they may prefer a higher availability tier (a stricter target at a higher price) for the peak period only. That turns a vague word into a priced choice.
Worked example of the final sentence
"The system shall be available 99.9% per calendar month, measured at the public endpoint, excluding up to 4 hours of notified maintenance; in a regional failure it shall be restored within 4 hours with no more than 15 minutes of data loss."
Trade-offs and pitfalls
- Promising 99.99% without the architecture (multi-zone, meaning copies in separate data centres within a region, or multi-region, meaning copies in separate geographic regions) and cost to support it.
- Leaving out how it is measured and what is excluded: those clauses decide disputes.
- Treating the SLA as the engineering goal: set the internal SLO tighter than the contractual SLA.
- Mixing up RTO and RPO: time to recover versus amount of data lost.
- What would change my numbers: the customer's actual cost of downtime and the price they will pay for each extra nine.
You get a one-page summary of a prospect's architecture: an on-prem database cluster behind a firewall, an API gateway in a private subnet, several third-party SaaS integrations and nightly batch jobs. Before recommending anything, what do you want to learn, what do you ask them to send, and where do you expect the integration risk to be?
Sample Answer
Direct answer
A one-page diagram shows boxes and arrows but not what flows through them, who owns them, or how they change. Before recommending anything I want to learn three things: what the solution must connect to and in which direction, what constraints the network and security rules impose, and how much of the current behaviour is undocumented. The five numbered questions below break those three themes into specifics (1 and 2 are about connections and direction, 3 and 5 about network and approvals, 4 and the data questions about undocumented behaviour). I would ask them to send interface details and policies, and I expect the integration risk to sit where the on-premises database, the private network and the nightly jobs meet the new system.
What I want to learn
- What business process this serves, and which data matters most.
- Which systems are the source of truth (the one system whose version of a record wins when systems disagree) for users, transactions and customer records.
- How data moves today: real time, batch, or by file, and with what volumes.
- Which of the third-party SaaS (software delivered as a hosted service) integrations are in scope, and who administers them.
- Who approves connections into and out of the network, and how long that takes.
What I ask them to send
- An interface inventory: for each link, source, target, protocol, data format, volume and schedule. A real row looks like: "Orders database (on-premises) to billing SaaS; REST API over HTTPS; JSON; about 40,000 records a night; 02:00 to 03:30; retries three times, then emails the operations mailbox." One row per arrow on the diagram is what turns the picture into facts.
- Firewall and network policy summary, including whether outbound internet access is allowed from the private subnet (a network segment not reachable from the public internet).
- Identity provider details (single sign-on, SSO, method and user directory).
- Batch job schedule with run times, durations and failure handling.
- Sample, anonymised records and data dictionaries (documents listing each field, its type and its meaning).
- Any existing security requirements for vendors.
Where I expect integration risk, by area
| Integration point | What it needs | Why it is risky here |
|---|---|---|
| User sync (keeping users and roles aligned between the customer's directory and our system) | Directory feed or standard provisioning (SCIM, System for Cross-domain Identity Management) | Directory is on-premises; a cloud service cannot reach it through the firewall without a connector (a small piece of software or a gateway that relays traffic between the two networks) |
| Transaction ingestion (ingestion means loading the customer's transactions into our system) | API calls or file drops, with ordering, retry and duplicate handling | Database behind the firewall means the customer must push out, or we need an agent (software installed inside their network) that only opens connections outward, so no inbound firewall hole is needed |
| Single sign-on (SSO) across cloud, customer relationship management (CRM) system and on-premises systems | A standards-based sign-in protocol (SAML or OpenID Connect, two common ways for an identity provider to tell an application who the user is) with their identity provider | Different systems may use different identity sources, so users can end up with several accounts |
| Nightly batch jobs | Agreement on the cut-off time and who is source of truth after the run | Real-time expectations clash with batch data that is up to a day old |
| Third-party SaaS | API limits, credentials and change notices | Rate limits (caps on how many API calls a vendor allows per minute or day) and vendor changes outside the customer's control |
My expectation, ranked
- Network connectivity to the on-premises database is the first blocker, because it needs a security approval that can take weeks. A typical network change request names the source and destination hosts, the port and protocol, the business reason, the data classification and an owner; it then waits in a queue for review by the network team, the security team and often a change board that meets weekly or fortnightly, which is where the weeks go.
- Data freshness: nightly batch versus any promise of live data.
- Identity: which directory is authoritative.
- SaaS dependencies, which I cannot control.
I commit to "network and data freshness first" because they gate everything else; a different customer with a modern network would flip this order.
Pitfall
Recommending a product on the basis of the diagram alone. The diagram says what exists; the risk lives in the interfaces it does not show.
What are some technical qualification questions you use early to decide whether an opportunity is a fit from an architecture perspective, for example across network, identity and data? What answers tell you it is a good fit and what answers signal a probable no-go?
Sample Answer
Direct answer
Early technical qualification means asking a short set of questions that reveal whether our product can be deployed into the customer's environment on their timeline. I use roughly three questions in each of three areas, network, identity and data, and for each one I know in advance which answers mean "good fit", which mean "fit with work" and which mean "probable no-go". The goal is to find a no-go in the first or second call, not at the proof of concept stage (a limited trial against agreed success criteria).
(Terms. Presales is the pre-contract technical function; SA means solutions architect, the presales engineer who speaks in the dialogue below. Qualification: deciding whether an opportunity is worth pursuing and on what terms. Architecture fit: whether the product works inside the customer's real technical environment. No-go: a deal I would stop pursuing or pause.)
The questions and how to read the answers
| Area | Question | Good fit | Fit with work | Probable no-go |
|---|---|---|---|---|
| Network | How will the workloads reach our service: public internet, private connection, or air-gapped (no internet at all)? | Public internet with allow-listing (the customer's firewall permits traffic to our addresses), or a private link we support | Private connectivity (a VPN, an encrypted tunnel over the internet, or a dedicated link, a physical private line) needing weeks of lead time | Fully air-gapped when we are a hosted service with no on-premises edition |
| Network | Are there firewall or proxy rules that block outbound calls to external endpoints? | A standard allow-list request | A proxy (an intermediary that all outbound traffic passes through) that inspects traffic and breaks our connections, for example by cutting long-lived ones | Policy bans any outbound traffic and no alternative is accepted |
| Network | Which regions or data centres must the service run in? | A region we operate in | A region on our roadmap with a date | A required in-country region we do not plan to have |
| Identity | How do users sign in today: SSO (single sign-on) through SAML or OIDC (OpenID Connect, both standard sign-in protocols)? | An identity provider we already support | Needs SCIM (automatic user provisioning) or custom roles | A proprietary directory with no standard protocol, and no budget to bridge it |
| Identity | Who is allowed to see what: role model, and is MFA (multi-factor authentication) mandatory? | Roles map onto ours | Fine-grained, attribute-based access (permissions decided by user properties such as department or location, not just a role) that we only partly support | A control we cannot enforce and that is a hard requirement |
| Identity | Do service accounts or API keys need rotation and central secret storage? | They use a secrets manager (a vault that stores and rotates passwords and keys) we integrate with | Custom rotation scripts | Short-lived credentials (that expire in minutes or hours) required, and we only offer static keys (long-lived secrets that keep working forever if leaked) |
| Data | What data will flow through, and how is it classified (public, internal, regulated)? | Non-sensitive or already-covered classes | Regulated data needing a contract addendum | Data class we are prohibited from handling (for example classified government data without authorization) |
| Data | How much data, how fast, and where does it live today? | Volume inside tested limits | Large one-off migration needing planning | Volume or latency beyond what the product has run |
| Data | Retention, deletion and who holds the encryption keys? | Our standard options | Customer-managed keys (the customer controls the encryption keys), which we support with setup | Customer must hold keys in their own hardware module (a dedicated tamper-resistant device that stores keys) and we cannot integrate |
The same follow-the-thread approach works in the other areas. Identity: "How do people sign in today?" "Our identity provider, with SSO." "Do you also need accounts removed automatically when someone leaves?" "Yes, within a day." That is an amber: it needs SCIM provisioning, so I note it as scope. Data: "What data would flow through?" "Customer records, some health data." "Then does a signed health-data agreement need to be in place before any pilot?" That is amber with a legal owner and a date.
That is nine questions in three groups of three. I would not ask them as a checklist. I ask the first question in each area, then follow the thread the answer opens.
How the conversation sounds
SA: "To make sure we deploy this where it works, how would your workloads reach an external service: out over the internet, or only through a private link?"
Customer: "Everything goes through a proxy, and we only allow a short list of approved vendors."
SA: "Understood. Is that proxy able to pass long-lived connections (one connection kept open for a long time), or does it cut them after a minute? Our agent keeps a connection open for streaming."
Customer: "I would have to check, I think it cuts them."
I have found the real issue: a proxy that cuts connections, plus a vendor approval list. That is "fit with work": I can propose a polling mode if we have one (our agent asks for updates every few seconds over short separate requests instead of holding one connection open, so a proxy that cuts long connections does no harm), and I note the vendor approval as a timeline risk. If we had no polling mode, I would mark it a probable no-go unless their network team can exempt us.
Deciding go, go with conditions, or no-go
- Go: no red answers, at most a couple of ambers with known workarounds and owners.
- Go with conditions: one or two ambers that decide the timeline, written down and agreed with the customer ("we proceed if your network team confirms the proxy exemption by the end of the month").
- No-go or pause: a red answer on a hard requirement the product does not meet and the customer will not relax. I say this to the account executive (the sales owner of the deal) with evidence, and I propose either revisiting when the roadmap closes the gap or declining politely. Continuing costs presales time and damages trust when the pilot fails.
Pitfalls
- Asking only about features. Most architecture no-gos come from network, identity and data, not features.
- Accepting "we will sort that out later" on a hard requirement. Record who will sort it out and by when.
- Treating a probable no-go as a negotiation. A "must" and a "nice to have" are different, so ask "if this were not possible, would the project stop?"
- Asking questions in a way that leads the customer toward the answer you want.
Walk me through how you map an enterprise customer's technology landscape layer by layer. At each layer, what do you ask, what integration points do you probe, and what signals tell you the implementation will be harder or costlier than it looks?
Sample Answer
Direct answer
I map an enterprise's technology landscape top-down in seven layers: business and users, applications, data, integration, identity and security, infrastructure and network, and operations. At each I ask what exists, who owns it, and how it connects to what I am proposing. The signal that an implementation will be harder than it looks is usually not a missing feature: it is an integration with no owner, a system nobody dares change, or an approval path nobody has described.
The seven layers
| Layer | What I ask | Integration points I probe | Signals it will be harder or costlier |
|---|---|---|---|
| 1. Business and users | Who uses the process, how many, which locations and peak periods | Where users start their day (portal, email, desktop apps) | Many user groups with different needs; no single process owner |
| 2. Applications | Which systems hold the relevant process today, which are packaged versus custom, which are being retired | APIs, file exports, screen-only legacy apps | Custom or end-of-life apps with no documentation or API |
| 3. Data | What data, where, how clean, who owns it, how large, how often it changes | Master data sources (the agreed single source for core records such as customers and products), migration paths, reporting feeds | Conflicting records between systems; no data owner; data quality unknown |
| 4. Integration | How systems talk today: APIs, message queues (a buffer where one system leaves messages for another to collect), file transfer, nightly batch (a bulk transfer run once overnight) | Each interface: direction, volume, format, error handling | Many point-to-point links (separate custom connections between each pair of systems); batch-only exchange when real time is needed |
| 5. Identity and security | Which identity provider, how access is granted, what policies apply | Single sign-on (SSO, one login across applications), directory sync (copying user accounts between directories), secrets management (secure storage of passwords and API keys) | No central identity provider; strict policies on external vendors |
| 6. Infrastructure and network | Where workloads run (on-premises, cloud), network zones (segments separated by rules), firewalls (filters that allow or block traffic) | Connectivity between zones and to outside services | Locked-down network; long firewall change lead times, or tight limits on egress (traffic leaving their network) |
| 7. Operations | Who runs it, monitoring, change process, support hours | Ticketing, monitoring and deployment tools | No dedicated team; long change windows (the only times changes are allowed); heavy change approval |
Seven layers is a mnemonic for me, not a rule: I merge layers for a small customer. If time is short I cover applications, integration, and identity and security first, because from experience (an ESTIMATE, not a measurement) undocumented integrations, access approvals and network lead times cause most late surprises.
Extension: on-premises Windows estate with strict data residency
Data residency (data must stay within a named country or region) adds questions at layers 3, 5 and 6:
- Which data fields are restricted, and does restriction apply to storage, processing or only transfer?
- Is the Windows estate joined to an on-premises Active Directory (Microsoft's directory service)? Do we need to sync identities without data leaving the region?
- Which regions or sites are acceptable? Is an in-country cloud region available?
- Are backups and support access also in scope? Remote vendor access can itself breach residency rules.
Inputs I need to scope a later proof of concept (POC, a time-boxed technical trial)
- One representative use case and its success criteria (the pass or fail results that would mean the trial worked).
- A current architecture diagram and an interface list.
- A sample or anonymised data set and its volume.
- A named technical contact and a test environment with access rules.
- Security and network approval paths, with lead times.
Reading the signals
Count three risk questions at every layer: Who owns this? How is it changed? What breaks if it is down? Where nobody can answer, assume the cost is higher than the customer thinks and say so in the estimate.
Mini example (illustrative): at layer 4 the customer says orders reach the ERP from the web shop through a nightly file copy, a script written years ago by someone who has left, and finance now wants same-day figures. Owner: nobody. Change: untested. If it breaks: orders stop. Signal: batch-only exchange when faster is needed, with no owner. In the estimate I add a new integration build plus time to reverse-engineer the script, say 3 to 4 weeks extra, and flag the risk in writing.
Pitfall
Interviewing only the sponsor. They know the business case but not the batch jobs.
A hospital network wants to use your ride-booking platform to move patients between facilities. During discovery, what non-functional requirements do you probe for, and which of them would most change the architecture you propose?
Sample Answer
Direct answer
For a hospital moving patients between facilities, I would probe the non-functional requirements (NFRs, the qualities a system must have beyond its features: speed, uptime, security, compliance) in six groups: availability, latency and dispatch timing, privacy and compliance, auditability, integration with hospital systems, and safety and accessibility rules. The ones that would most change the architecture are privacy and compliance (it forces how data is stored, isolated and logged) and availability (it forces redundancy and failover), followed by integration with clinical systems. A seventh point sits outside the six probing groups because it comes from the product's own behaviour rather than from a hospital requirement: pricing and fairness rules (below).
What I probe, and why each matters
| NFR area | Questions I ask | Why it matters |
|---|---|---|
| Availability | "What happens if a transfer cannot be booked for 30 minutes? Is there a manual fallback?" | Sets uptime target and failover design |
| Latency and timing | "How fast must a dispatcher (the person or system that picks and sends the vehicle) see a vehicle assigned? How fresh must the ETA be?" | Drives caching, real-time update design |
| Privacy and compliance | "Will patient health information (PHI) be in the booking? Which regulations apply (for example HIPAA, the US Health Insurance Portability and Accountability Act)? Will we need a signed business associate agreement (BAA, the contract in which a vendor handling patient data promises to protect it as HIPAA requires)?" | Changes data model, encryption, tenant isolation (keeping one customer's data and workloads apart from every other customer's), which cloud services are allowed |
| Auditability | "Who must be able to see who accessed or changed a booking, and for how long are logs kept?" | Immutable audit logs (records that cannot be altered once written), retention |
| Integration | "Which systems must exchange data: patient scheduling, electronic health records (EHR)? Which standards, such as HL7 (Health Level Seven) or FHIR (Fast Healthcare Interoperability Resources), the common message and API formats hospital systems use to exchange patient data?" | Adds an integration layer and mapping work |
| Safety and accessibility | "Wheelchair or stretcher vehicles? Trained or vetted drivers? Escort rules?" | Changes the matching logic, not just the infrastructure |
That is six rows because the question deserves breadth; in a real call I would prioritise by asking which two the hospital would fire a vendor over.
Which ones change the architecture most
- Privacy and compliance. A consumer ride-booking app (riders request a vehicle, dispatch assigns one, and an ETA, the estimated time of arrival, is shown) mixes data from many riders. A hospital deployment may need PHI kept apart from general data, encrypted with customer-controlled keys (encryption keys the hospital holds, so we cannot read the data without them), restricted by region, and logged on every access. That can mean a separate tenant or deployment, not a configuration flag. Concretely: the consumer app keeps all riders in one shared database, whereas the hospital would get its own database and encryption key in an approved region, so its patient records never sit beside other customers' data.
- Availability. A missed patient transfer has clinical consequences, so single-region, best-effort uptime is not enough. That pushes toward multi-zone deployment (running copies in several separate data centres in one region so losing one does not stop service), failover (automatic switch to a standby copy), a tested fallback process and a stricter service-level agreement (SLA, the contractual uptime and response promise).
- Pricing and fairness rules (not one of the six groups; I raise it because our own product behaves differently from what the hospital needs). Surge pricing (raising fares when demand is high) is acceptable for consumer rides and plainly unacceptable for patient transport, so the pricing service must be switchable per customer.
Linking each business objective to one technical metric
The hospital's goals are business sentences. I convert each into one number I can test, using the surge-pricing example for contrast:
| Business objective | Single technical metric |
|---|---|
| Staff trust the price they see before booking | Price displayed within a latency target (for example, the 95th percentile, P95, under 1 second) |
| Prices never punish patients | Fixed-fare rule enforced: 0 bookings priced above the contracted rate (an invariant: a rule that must always be true, and a test can check it) |
| Vehicle status is trustworthy | Location and status updates delivered within an agreed update SLA (for example, within 10 seconds) |
| Regulators can audit us | 100% of PHI access events present in the audit log, retained for the period the contract states |
The figures above are illustrative: I ask the customer for their real thresholds.
Pitfalls
- Accepting "high availability" as a requirement. It is a wish until it has a number and a consequence.
- Treating compliance as paperwork at the end. It decides the architecture at the start.
Unlock Full Question Bank
Get access to all 14 Technical Discovery & Needs Qualification interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.