Proof of Concept and Demonstrations Questions
Proving technical fit through demos, proofs of concept, pilots, and evaluations. Covers scoping and running a POC or pilot, defining success criteria and acceptance tests, and handing a successful result over to delivery. Also covers designing and delivering demonstrations for different audiences, preparing demo environments, realistic data and sandboxes, recovering when a live demo fails or draws hard questions, and validating performance, integration, and security requirements. Focuses on the demonstration and validation craft.
During a proof of concept the customer reports inconsistent test results and asks you for help. How do you work out whether the cause is the environment, the data or the product, and how do you communicate what you find?
Sample Answer
Direct answer
I treat it as a controlled experiment: change one thing at a time until the variation follows one of three suspects. First I get the facts (what varies, by how much, since when), then I compare environment, data and product separately, then I report what I found in plain language with evidence. The 2 x 2 swap is the workhorse: run their data on our reference environment, and our reference data on their environment.
Step 1: pin down the symptom
Ask for 3 or more runs with timestamps, the exact steps and the raw output. Quantify: "results differ by 20% between runs" is a different problem from "run 3 crashed". Ask what changed since the last good run (code, config, data load, other tenants on the system, tenants being other customers' workloads sharing the same machines). Do not guess before you have numbers.
Step 2: separate the three suspects
| Suspect | Typical evidence | Quick check |
|---|---|---|
| Environment | Variation tracks time of day, noisy neighbours (other workloads on shared hardware that steal capacity), undersized or throttled instances (slowed down deliberately once they use up their allowance), different versions or config | Diff the environment manifest (a list of exactly what is deployed: versions, instance sizes, config checksum, network path) against the agreed spec; a checksum is a short fingerprint of a file, so a changed file has a different one |
| Data | Differences vary with which records are loaded, skew (a lopsided mix, such as a few huge records or one value dominating), duplicates, stale or partially loaded data | Compare row counts and checksums between runs; rerun on a fixed seeded dataset (generated from a fixed starting number so it is identical each time) |
| Product | Same data, same environment, still varies; a fault I can reproduce on my own reference environment | Reproduce on our reference build with a fixed dataset and file it with engineering |
The swap test makes this decisive:
| Our environment | Customer environment | |
|---|---|---|
| Our data | Baseline, should be stable | Isolates the environment |
| Customer data | Isolates the data | The reported symptom |
If the problem follows the customer's data into our environment, it is data. If it follows their environment with our data, it is environment. If both stay fine and only the original combination misbehaves, suspect an interaction (the two only misbehave together) and look at configuration: for example a setting tuned for small data, such as a memory limit or a batch size, that only hurts at the customer's data volume. If it reproduces everywhere, it is the product.
Step 3: communicate
- Report on the same day the finding is confirmed, even if the cause is not: "Here is what we know, what we ruled out, and what we do next."
- Use evidence the customer can verify (their own logs, side-by-side result tables) and do not assign blame early. Say "the variation follows the dataset" rather than "your data is wrong".
- If it is our product, say so plainly, give the workaround, the fix owner and a date range. Credibility from owning a defect is worth more than the defect costs.
- Record the finding in the POC status document and say whether the success criteria or timeline need adjusting.
Worked example
A customer sees query times of 1.2 s, 1.9 s and 1.3 s. Their environment manifest shows a smaller instance type than agreed. That type is burstable: it earns CPU credits while idle and spends them under load, and once the credits run out the cloud provider slows the CPU to a low baseline, so a run right after a burst is slow and one after an idle spell is fast, which explains times drifting between 1.2 s and 1.9 s. Filling in the swap table (illustrative numbers): our data on our environment, 0.7 s on all three runs (baseline, stable); customer data on our environment, 0.7 s on all three runs (so their data is not the cause); our seeded data on their environment, 1.2 s, 1.9 s and 1.3 s (the variation appears, so it is the environment); customer data on their environment, 1.2 s, 1.9 s and 1.3 s (the reported symptom). The variation follows the environment, not the data, so the environment is the cause. Fix: move to the agreed instance size, rerun, and note the result in the status report. A sample message to the customer (illustrative): "Your three runs varied from 1.2 to 1.9 seconds. Your servers are the smaller size, not the one in the plan, and they slow down after bursts of work. The same data on our reference environment runs at 0.7 seconds every time. Please move to the agreed size by Wednesday; we will rerun together on Thursday and update the status report. Nothing so far changes the success criteria."
Pitfalls
Changing two variables at once, trusting a one-off run, defending the product before the evidence exists, and burying a real product bug because the POC is close to its end date.
A prospect wants to see your product working against their own SaaS data source. How do you set up a customer-specific sandbox for it, and how do you handle credentials, access and clean-up?
Sample Answer
Direct answer
I would build an isolated, time-limited sandbox in my own cloud account, connect it to the prospect's test or sandbox copy of their SaaS (software as a service) system using a narrowly scoped read-only credential that they create, pull only the data we agreed, and delete everything with a written confirmation at the end. The order matters: agree the rules in writing first, build second, clean up on a fixed date.
Setup, step by step
- Agree scope in writing. Which source, which objects (for example, only Accounts and Cases), which fields are off limits, how long it runs, who has access, and what happens to the data afterwards. A short data-handling note or addendum to the non-disclosure agreement (NDA, the contract in which both sides promise to keep shared information confidential).
- Use their non-production tenant (a tenant is the customer's own separate instance of the SaaS product; non-production means a test copy, not the one their business runs on). Ask for a developer or sandbox environment of the SaaS tool (most vendors offer one) populated with test or masked data (real values replaced by realistic fakes), not production. If only production exists, use a subset filtered by them.
- Isolate per prospect. One separate cloud account or project per prospect, defined as infrastructure as code (configuration in files, so it can be created and destroyed identically). Tag every resource with the prospect, the owner and an expiry date. What that looks like in Terraform, a common infrastructure-as-code tool (the names are placeholders):
resource "aws_s3_bucket" "sandbox_data" {
bucket = "acme-demo-sandbox"
tags = {
prospect = "acme"
owner = "jsmith"
expires = "2026-11-14"
}
}
A nightly job can list every resource whose expires tag is in the past, and clean-up later runs terraform destroy on the same files.
4. Ingest minimally. Ingest (copy in) only agreed objects, with filters (date range, record count), and mask or drop sensitive fields as they arrive. Example: the email maria.lopez@acme.com is stored as user-8f3a@example.com, and the phone number is dropped.
5. Show the result. The demo runs on their data in the sandbox, and I note what I saw so I can describe it without re-accessing it.
Credentials and access
| Concern | How I handle it |
|---|---|
| Who creates the credential | The prospect's administrator, so I never see a password |
| Type | An OAuth (an authorisation protocol where the user grants limited access without sharing a password) app or an API token, read-only, with the narrowest scopes (scopes are the individual permissions on a token; for a CRM, something like cases.read and accounts.read and nothing that can write or delete; the names vary by vendor) |
| Storage | In a secrets manager (a managed vault service for passwords and keys, such as AWS Secrets Manager) in the sandbox account, never in code, chat or email |
| Rotation and lifetime | Set to expire at the end date; a refresh token (a long-lived renewal credential) is revoked at clean-up |
| Who on my side | Named people only, through single sign-on (one company login) with multi-factor sign-in; access logged |
| Network | The prospect allow-lists (adds to their list of approved addresses) my sandbox's outbound IP address, the fixed address its traffic leaves from, if their SaaS requires it |
Clean-up
Agree the date at the start (for example, 14 days after the last working session). On the day:
- The prospect revokes the token or app authorisation on their side (this is the step that matters most, since it works even if I forget).
- I destroy the infrastructure with the same code that built it, which removes the data, snapshots (point-in-time copies of disks or databases) and backups.
- I check no copies remain (logs, exports, local files).
- I send a short written confirmation listing what was deleted and when.
Worked example
Prospect uses a CRM (customer relationship management) system and wants to see our analytics on their support cases. We agree: Cases and Accounts only, last 12 months, email and phone fields masked, sandbox lives 14 days, one named engineer plus me. Their admin creates a read-only app and shares the token through the agreed secure channel. I store it in the secrets manager, load a filtered case sample, run the demo, and on day 14 they revoke the app while I destroy the account and send the deletion note.
Trade-offs and pitfalls
- Using production data is faster and more convincing but raises privacy and contract risk; I only do it with a subset, written approval and their legal sign-off.
- A long-running sandbox that nobody owns is the common failure. The expiry tag and the calendar reminder are the safeguards.
- Do not copy their data to a laptop for convenience.
A prospect wants proof that your product works with their identity provider within 48 hours, as a proof of concept. How do you run it, what do you need from them, how do you keep it safe, and how do you present the limits if full production integration is not possible?
Sample Answer
Direct answer
I would say yes to a 48-hour proof of concept (POC, a time-boxed experiment that tests whether the product works in the customer's setting) only for a narrow, written scope: single sign-on (SSO) for a handful of test users against a non-production test tenant (a separate, throwaway instance of their identity system, not the real one) of their identity provider (IdP, the system that holds user accounts and vouches for them, for example Okta or Microsoft Entra ID). I would ask them for a test tenant, a named admin, the IdP's metadata and a few synthetic users, I would never take a production password or secret in chat or email, and I would close with a readout (a short results meeting and report) that separates "proved" from "still to prove in production".
How I run the 48 hours
| Window | What happens | Output |
|---|---|---|
| Hour 0 to 2 | Kickoff call: confirm protocol (SAML, Security Assertion Markup Language, or OIDC, OpenID Connect), the three scenarios to prove, and who is on call from their side | One-page scope, agreed in writing |
| Hour 2 to 24 | I configure our side from their metadata; their admin registers our application in their test IdP | First successful login |
| Hour 24 to 40 | Test matrix: login, logout, wrong user denied, group-to-role mapping, session timeout | Pass/fail log with screenshots |
| Hour 40 to 48 | Readout: results, limits, and the production path | Short report and next-step plan |
Worked test case, with fictional values: user ana@test-corp.example in group poc-admin logs in and must land in our product with the Admin role; ben@test-corp.example in poc-viewer must land as Viewer; cho@test-corp.example, in neither group, must be refused with an access-denied page. Scope I would write down: (1) a user from the test IdP logs in with SSO, (2) a user outside the allowed group is refused, (3) a group in the IdP maps to a role in our product. If they also want automatic user provisioning (SCIM, System for Cross-domain Identity Management), I would list it as out of scope unless the first three pass early.
What I need from them (five items)
- A test or development IdP tenant with an admin who can register an application. Not production.
- The IdP metadata: for SAML the metadata XML or URL (a file describing their login endpoint, signing certificate and entity ID), for OIDC the issuer (discovery) URL (an address that publishes the same details).
- Two or three synthetic test users in two groups (for example
poc-admin,poc-viewer), with MFA (multi-factor authentication) behaviour stated so I know what to expect. - The attributes they want sent (email, name, group) and which one is the unique user identifier.
- A named contact reachable for the whole window, plus anyone who must approve a change on their side (for instance a security team).
I send them from my side: our SAML entity ID (the unique name our application goes by) and callback (assertion consumer service) URL (the address their system posts the signed login result back to), or the OIDC redirect URI (the same idea for OIDC), so they configure their end and I never need access to their admin console.
How I keep it safe
- Test tenant and synthetic users only. No real employee data, no production IdP.
- No credentials by email or chat. If an OIDC client secret must cross, it goes through a one-time secret link or their vault, is scoped to the test app, and is rotated or deleted at the end.
- Least privilege: their admin does the IdP changes while screen-sharing; I get no standing admin rights.
- Our side is a throwaway environment for this prospect, with a fixed expiry and teardown at hour 48 or when they say stop.
- Log what was done (who configured what, when) so the security reviewer can inspect it.
The directory (LDAP / Active Directory) variant
LDAP (Lightweight Directory Access Protocol) and Active Directory (AD, Microsoft's directory service) differ because the product usually needs a bind account (a service login to search the directory). My approach: do not ask for a real bind password. Options in order of preference: (a) a connector or agent running inside their network that they configure themselves, so the secret never leaves; (b) a read-only service account limited to a test organisational unit (OU, a folder-like container in the directory) with synthetic users, connected over LDAPS (LDAP over TLS), the secret shared as above; (c) if neither is possible in 48 hours, I stand up a sample directory (for example an OpenLDAP container with fake users and their real attribute naming) and prove the mapping logic there. Minimal configuration: server address, search base (where in the directory tree to look), user filter (which entries count as users), attribute map and group map. For example the customer enters server ldaps://ldap.test-corp.example:636, search base ou=poc-users,dc=test-corp,dc=example, user filter (mail={email}), and maps attribute memberOf to our roles. I would be explicit that (c) proves our logic but not their network path.
Presenting the limits
I would use three columns in the readout, spoken in plain words:
| Proved in 48 hours | Not proved | What production needs |
|---|---|---|
| Login, deny, group-to-role mapping on test tenant | Their production IdP policies (conditional access, meaning rules such as 'only from managed laptops'; signing certificates rotation, meaning the periodic replacement of the certificate that signs logins), scale of real user population, SCIM | Change request, production tenant app registration, certificate expiry plan, security review, pilot with real users |
Example lines: "What you saw is a working login against your test tenant. It does not tell you how this behaves with your conditional-access rules or 8,000 real users, and I would not want you to read it that way. Here is the two-week path that does."
Trade-offs and pitfalls
- Saying yes to everything in 48 hours produces a half-proved demo that later reads as a failure. A smaller, fully passing scope beats a wide one.
- The usual delay is on their side (waiting for an admin), so I name the admin and a fallback before hour 0, and the clock starts when the test tenant is ready, not at the email.
- If a call is needed to unblock a stuck assertion, check clock skew, the audience/entity ID and the callback URL first, because mismatches in those three are among the most common causes of a failed first login in practice (an experience-based observation, not a measured share). Clock skew: login assertions are valid only for a few minutes, so a server clock that is 10 minutes off makes a fresh assertion look expired. Audience/entity ID: the assertion names who it is for; if their system says
https://app.vendor.example/samland we registeredhttps://app.vendor.example/saml/, the trailing slash alone makes us reject it. Callback URL: if the IdP sends the user to a different address than the one we registered, it refuses or the browser lands on an error.
A buyer is nervous about delivery risk and the learning curve of your product, and suggests a pilot. How would you design it so it genuinely settles those concerns, and what would a clean handover into production look like?
Sample Answer
Direct answer
I design the pilot so that each fear gets its own test with a pass mark. A pilot differs from a proof of concept: a POC checks that the product can work, a pilot checks that it works in the buyer's real operating conditions with real users. For delivery risk, the pilot delivers one real piece of the production integration on the buyer's own infrastructure, with their staff in the loop. For the learning curve, it measures how quickly their own people become productive without our help. Handover into production is planned from day one: what runs in the pilot is built to be kept, with owners, runbooks and a decision date.
Mapping each concern to a test
| Buyer concern | Pilot test | Pass mark (illustrative) |
|---|---|---|
| Delivery risk | Deliver one real integration into their pre-production environment (a realistic copy of production used for testing), with a rollback (reverting to the previous working state) rehearsed | Completed by week 4; rollback restores the previous state in a rehearsal |
| Learning curve | Two of their own admins complete five defined tasks with the documentation only | Both finish all five tasks by week 3 with nobody on our side doing a step for them (unassisted), and no more than two logged support requests each |
| Disruption | Run alongside the current process (shadow mode: the new system processes the same inputs but its outputs are not yet used) | No change to live operations; results compared daily |
| Operational fit | Monitoring and alerts wired into their tools | Alerts reach their on-call channel in a test |
How a pass mark is counted: every request for help is logged in a shared sheet with the date, the task, and whether the documentation was missing or unclear. Say admin A finishes all five tasks with 1 request, and admin B finishes all five tasks, with nobody doing a step for them, but logs 3 requests (all about missing documentation). A passes. B fails the "no more than two requests" mark although the tasks got done, and the fix is to patch the documentation and have B rerun the task, not to relax the mark.
Design in six steps
- Scope: one team, one workflow, one integration. The smallest slice that is still production-shaped.
- Minimal access: read-only first, then scoped write access, with named service accounts (logins for software rather than people) that can be revoked.
- Responsibilities: a written split of who does what on both sides, including the buyer's change-approval steps (their formal process for approving changes to live systems).
- Safety: shadow mode first, then a controlled cutover (the switch from the old path to the new one) for a small user group, with a documented rollback.
- Measures: weekly review of agreed KPIs (key performance indicators): error rate, time to complete the task, support requests, user satisfaction.
- Sign-offs: named stakeholders approve at each gate (a checkpoint before the next stage): security, operations, business owner.
If the buyer says they cannot afford any disruption
Show a six-week validation with three stages: weeks 1 and 2 shadow mode with no effect on live work, weeks 3 and 4 a small group of users on the new path with the old path still running, weeks 5 and 6 a go or no-go review (a formal decision to proceed or stop). The rollback strategy is rehearsed in week 2 and again before cutover. If their evaluation runs six months, I keep the same gates but add monitored expansion stages from week 7 (for example 5% of users in months 2 and 3, 25% in months 4 and 5, everyone in month 6, each starting once the six-week validation is signed off, each stage opening only if the previous one's error rate and support-ticket marks held), and I define exit criteria (the conditions that must hold to move on or to stop) for each stage so either side can stop at a gate without penalty.
What a clean handover looks like
- A decision before the pilot starts: the pilot environment becomes production, or it is rebuilt. I prefer the former for the integration and configuration, because it avoids repeating the work.
- Artefacts delivered: configuration stored as code (settings kept in version-controlled files so the environment can be rebuilt exactly), a runbook (step-by-step instructions for operating and fixing the system), an architecture diagram, an ownership table (who is paged for what), a support route and agreed response times. A runbook entry reads: "Symptom: nightly ingest job fails. Check: job log for authentication errors. Fix: rotate the service account key, rerun. Escalate to: our support desk if it fails twice." An ownership table row reads: "Ingest job failure | buyer platform on-call | 15-minute response"; "Product defect | our support desk | next business day".
- Commercial: pilot terms convert into the main agreement without renegotiation if the exit criteria are met.
- People: the buyer's staff run the first production release while I shadow, not the reverse.
Trade-offs and pitfalls
A realistic pilot costs the buyer's time and some of mine. A pilot with no pass marks only confirms what everyone already believed. A pilot built as a throwaway makes the handover a second project, which is exactly the delivery risk the buyer feared. If their main worry is learning curve, I weight the user tests over the integration test, and the other way round.
What is the difference between a proof of concept, a proof of value, a pilot and a production deployment? For each, what is it meant to prove and who signs off?
Sample Answer
Direct answer
They are four steps of increasing realism and commitment. A proof of concept (POC) proves something is technically possible in a controlled setting. A proof of value (POV) proves it is worth it: the customer's own scenario is run against a measured baseline (today's numbers), and the result is stated in business terms such as hours saved, cost avoided or revenue gained, in a form finance accepts. A pilot proves it works for real users in real conditions at limited scale. Production is full deployment under normal operations and support. Companies use these words loosely, so I always confirm what the customer means before agreeing scope. (SA means solutions architect, the pre-sales technical lead.)
Comparison
| Stage | Meant to prove | Typical setup | Who signs off |
|---|---|---|---|
| POC | Technical feasibility: does it work with our systems? | Sandbox, sample or synthetic data, days to weeks | Technical lead or architect |
| POV | Business value: does it move the metric that matters? | Customer's own scenario and data, a measured baseline, weeks | Business owner or sponsor, with finance for numbers |
| Pilot | Operational fit: does it work with real users, support and processes? | Limited users or one team, real data, under a short agreement, 1 to 3 months | Business owner, plus security/IT and sometimes legal or procurement |
| Production | Sustained service at scale under SLA (service-level agreement) | Full rollout, support, monitoring, full contract | Executive sponsor, procurement and IT operations; the contract is signed |
Contractual and operational differences: scoped POC vs pilot
| Aspect | Scoped POC | Pilot |
|---|---|---|
| Paper | Often a short evaluation or trial letter, sometimes unpaid | Pilot agreement, often paid or a credit against the purchase |
| Data | Synthetic or masked | Real data, needs security review and data handling terms |
| Acceptance criteria | Technical pass/fail tests | Technical plus usage and business metrics |
| Success metrics | Example: 95% of sample test cases pass | Example: 60 of 80 invited users active weekly and task time down against a baseline |
| Responsibilities | Mostly our SA and their architect | Customer runs it with real users, we provide support, training and an escalation path |
| Outcome | Decision to proceed to a pilot or value test | Decision to buy, with a rollout plan |
Worked example
A bank evaluates a document-processing tool. POC: 100 sample statements extracted with 95% field accuracy in a sandbox, signed by the architect. POV: on 1,000 real documents (masked, meaning names and account numbers replaced with fake values), manual handling time falls from 12 minutes to 4, a saving of 8 minutes per document. Finance converts it (illustrative inputs): 120,000 statements a year x 8 minutes = 960,000 minutes = 16,000 hours; at a loaded labour cost of $35 an hour that is 16,000 x $35 = $560,000 a year, to be set against the licence price. Pilot: one operations team of 30 uses it for two months under a pilot agreement (a short contract that sets scope, duration, data handling and what happens at the end). Production: contract, SLA and rollout across all teams.
Pitfalls
Calling a free long trial a "POC" with no criteria, so it never ends; skipping the POV and then being unable to justify the price; or running a pilot without acceptance criteria and sign-off, so the customer drifts on free usage.
Unlock Full Question Bank
Get access to all 25 Proof of Concept and Demonstrations interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.