Survey or Questionnaire
A survey, or questionnaire, gathers information from a group of people in a structured way over a short period. The questions are administered in writing, online, by telephone or in person, and the respondent answers either by choosing among predefined options, the closed question, or by writing freely, the open question. The technique runs in three stages: prepare the instrument and the sample, distribute and collect, then process the responses. What it offers is reach, contacting three hundred people costs almost as much as contacting thirty, and closed questions produce countable data that lends itself to statistical analysis. What it demands in return sits in two numbers that most surveys neglect: the response rate, which says how many of the people targeted actually answered, and the margin of error, which says how precisely the sample result holds for the population. A survey without these two numbers produces percentages that only have the appearance of a measurement.
Goal
The survey elicits business information from a group, through one set of identical questions put to each person, and it delivers a result that can be quantified: "fifty-eight percent report being satisfied, give or take three points". Part of what an organisation needs to know is spread across a great many heads: what a whole workforce thinks, what hundreds of members of an association want, how a geographically dispersed customer base judges a service. Interviewing these people one by one would cost time nobody has, and would return impressions that cannot be counted or compared.
Four properties give it its value, and BABOK lists them as the strengths of the technique. Reach, first: the survey touches a wider audience than the interview at an almost negligible marginal cost per additional respondent. Low burden next, on the respondent and on the analyst alike, once the instrument is built. Geographic span, which reaches stakeholders scattered across several sites effortlessly. And above all quantification: closed questions produce numerical data that can be worked statistically, and open questions surface observations the other techniques miss, because they did not know to ask for them.
The deliverable is a set of structured responses. Closed questions yield proportions with their margin of error, distributions and cross-tabulations; open questions yield a body of text that is coded into themes and categories. The value of this deliverable hangs entirely on two decisions taken before the first send: how the sample was drawn, and what response rate was reached. A percentage drawn from a biased sample or a derisory response rate measures nothing.
The survey is chosen on a trade-off: reach against depth. It buys coverage and the count, and it pays for them in shallowness and non-response bias. Where the interview follows up, challenges and tracks an answer in real time, the survey records only what the respondent chose to write in the box. The most productive combination fits in one sentence: three interviews say which questions to put to the other three hundred people and in what words.
Usage
When to use it
- Large or dispersed population: a workforce, a membership or a customer base that the interview cannot reach one by one has to be covered.
- Numbers are needed: proportions, distributions and statistical analysis are wanted.
- The range of answers is known: closed questions fit when it is already clear which options the respondent can choose.
- Low cost and burden: reaching hundreds of people costs almost the price of reaching dozens, for little time asked of each.
- Repeated measurement over time: the same instrument measures a satisfaction or an opinion year on year and makes the trend comparable.
- Following interviews: size across three hundred people what three interviews surfaced, to learn whether the finding holds beyond the three.
- Feedback during a large change: gather by functional unit the mood of a rollout, to act before opposition sets in.
When not to use it
- The "why" behind the number: the survey measures the scale of the phenomenon; the cause comes from an interview or a focus group.
- The range of answers is still unknown: closed questions harvest only what was foreseen, the interview surfaces the unexpected and then dictates the options.
- The information is already on record: it lives in the files, run a document analysis.
Description
Two question types, four closed forms
BABOK sets out two question types. The closed question makes the respondent choose among predefined options. It is easier to analyse, attaches to numerical values and feeds statistical processing, and it is used when the range of possible answers is well understood. The open question calls for a written answer. It yields detail and unforeseen responses, and it is paid for in analysis: free text is slow and subjective to categorise. It is used when the issues are known but the range of answers is not.
The closed family comes in four forms, and the choice of form decides what can be done with the answers.
| Closed form | What the respondent does | What you get | When to use it |
|---|---|---|---|
| Dichotomous (yes/no) | Chooses between two options | A binary proportion, immediate to aggregate | Filter, confirm a fact, route the rest of the questionnaire |
| Multiple choice | Selects one or more entries from a list | A distribution over known categories | When the possible options are already well mapped |
| Ranking | Orders options by preference | A priority order between items | Prioritise competing needs or criteria |
| Rating scale (Likert) | Places agreement on an ordered scale | An ordinal measure, summarised as a positive grouping | Measure an attitude, a satisfaction, a degree of agreement |
Wording decides the measurement. The Pew Research Center's work shows that a minor change of phrasing, order or answer options routinely shifts the measured opinion by several percentage points and sometimes by more than ten. Three practical consequences follow. A leading question, "would you not agree that...", plants the answer it claims to measure; ask for the real opinion, "how do you rate...". A double-barrelled question, which asks two things at once, "are pay and prospects satisfactory", cannot receive a clean answer and splits in two. And an ordinal scale is presented in its order, because the order is the information. Open questions, finally, are used sparingly: their per-item non-response rate sits around eighteen percent, against one to two percent for a closed one, and a questionnaire that overuses them empties out along the way.
The three stages: prepare, distribute, document
Prepare. Define the objective, then the target population with its size and its variations of culture, language and location. Choose the balance of closed and open questions, and decide on the sample: survey the whole population or draw a subset by a sampling method. Settle the modes of distribution and collection, a target response rate and a schedule, and decide whether interviews should be paired with the survey, which does not offer the depth interviews do. Write the questions from the objective, each serving a precise purpose. Then test the questionnaire on a few respondents: this field test surfaces ambiguous questions, broken logic jumps and mis-calibrated scales before the send, while they can still be fixed. A badly designed instrument is not repaired after the fact: it has to be rebuilt and put back in the field.
Distribute. Communicate the survey's objective, the use that will be made of the results and the confidentiality or anonymity arrangements in place. These three messages decide the response rate and the frankness of the answers. Choose the channel by urgency, the level of security required and geographic spread. Anonymity, once promised, must be real down to the way results are reported, failing which respondents sense it and fall silent.
Document. Gather and summarise the responses, draw out the themes that emerge from the open questions, formulate the coding categories and break the data into measurable increments. This is where closed questions become proportions and free text becomes a set of counted themes. Data mining takes over when the volume of closed responses warrants a deeper pattern analysis than simple tabulation.
The numbers that keep a survey honest
Three numbers turn a heap of responses into a defensible measurement, and confusing them is the source of almost every error in the discipline.
The response rate is the number of usable responses divided by the number of eligible people who actually received the questionnaire. The denominator is the list of reachable, in-scope recipients: an invalid address, an out-of-scope recipient leave the denominator. This is AAPOR's point: a response rate is computed on the eligible sample.
The margin of error of a proportion is z·√(p(1−p)/n), where p is the observed proportion and n the sample size. At the 95% confidence level z is 1.96; use 1.645 at 90% and 2.576 at 99%. When p is unknown at the time of designing the survey, take the worst case, p equals 0.5, which maximises p(1−p).
The finite-population correction comes in when the sample represents an appreciable share of the population, in practice beyond about 5%. The margin is then multiplied by √((N−n)/(N−1)), where N is the population and n the sample. When N is much larger than n, this factor is roughly 1 and can be set aside; when n is a large fraction of N, it tightens the interval markedly. The rule is to apply it fully or declare it negligible.
To size the sample in advance, Cochran's formula gives the size n₀ = z²·p(1−p)/e² needed for a target margin e, then adjusts it to a finite population by n = n₀/(1 + (n₀−1)/N). This is the calculation that answers the question "how many responses do I need", and it is done before the send.
From here, the central trap: a high response rate does not on its own give a small margin of error, it is n that governs it. Ninety percent of responses on a list of thirty people is twenty-seven responses, with the margin of error of twenty-seven responses. The response rate protects against non-response bias, the sample size determines precision, and the two numbers answer two different questions that must not be mixed.
One last factor governs both the response rate and the margin of error: the way the sample is drawn. Only a probability sample, where every unit of the population has a known and non-zero chance of being selected, allows the result to be generalised to the population with a quantifiable margin of error. A non-probability sample, a self-selection panel or a convenience sample, yields only a participation rate, which AAPOR distinguishes from the response rate, and it carries a selection bias that a large n does not correct. A thousand responses drawn from volunteers alone measure the opinion of volunteers, and the margin of error computed on them is a mathematical ornament laid over a sample that represents no one known.
AI considerations
AI acts at both ends of the technique, design and processing. In design, a language model produces a first set of questions from the objective, and it excels at the mechanical part of the work: turning an intent into a closed question, proposing the options of a multiple choice and above all flagging the leading phrasings and double-barrelled questions a writer lets through without seeing them. Simulating a field test, by having the model answer as different respondent profiles, surfaces some ambiguities before the send. In processing, AI earns its place on the open questions: grouping hundreds of free-text responses, proposing coding categories and summarising the themes is the step BABOK calls "formulate categories for encoding the data", and it is work where the machine handles in minutes what took days. This is also the analytical angle the Guide to Business Data Analytics takes: structure the questions around a hypothesis to test and track the responses as they come.
The limits are sharp. AI does not certify the validity of the sample: a biased sampling frame stays biased whatever the sophistication of the analysis that follows, and no model rescues a badly drawn population. On free text, it fabricates plausible themes from thin material, and an invented theme looks like a real one: the analyst checks every category against the verbatims that populate it. The decision on statistical significance and on the response rate stays a human judgement; the margin-of-error calculation, for its part, obeys a formula. Finally, the anonymity and sensitivity of raw responses fall to human governance, which the Guide to Business Data Analytics underlines: handing a corpus of named responses to an external service communicates personal data, and the decision is made on data-protection grounds.
Examples
A services company based in Fribourg runs its annual satisfaction survey among its 480 employees. It surveys them all, so this is a census. The questionnaire is anonymous, it mixes four question forms and takes five minutes to answer.
| No. | Question | Form | Options |
|---|---|---|---|
| Q1 | Overall, I am satisfied working at the company. | Likert scale | Very satisfied · Fairly satisfied · Neutral · Fairly dissatisfied · Very dissatisfied |
| Q2 | Would you recommend the company to someone close to you looking for a job? | Dichotomous | Yes · No |
| Q3 | Which area should be improved first? | Multiple choice (one answer) | Pay · Workload · Internal communication · Career prospects · Working conditions |
| Q4 | What would make you recommend the company more readily as an employer? | Open | free field |
312 employees answer usably. The response rate is therefore 312 / 480, or 65.0%. On the headline question, 58% report being satisfied or very satisfied, that is the proportion p = 0.58 over n = 312 responses. The margin of error at 95% is first 1.96 × √(0.58 × 0.42 / 312), or ±5.5%. Since the 312 respondents represent 65% of the population of 480, the finite-population correction applies: it is √((480 − 312) / (480 − 1)), or 0.592, and it brings the margin down to 5.5% × 0.592, or ±3.2%. The honest result therefore reads "58% satisfied, interval from 54.8% to 61.2% at 95% confidence".
The lesson of the case turns on the exact role of each number. The response rate of 65% is good and protects against non-response bias, but the precision comes from the 312 responses, tightened by the finite-population correction, which produce the ±3.2%. A sizing calculation confirms it: to target ±5% at 95% with an unknown proportion (p = 0.5), Cochran's formula gives n₀ = 1.96² × 0.25 / 0.05², or 384, then n = 384 / (1 + 383 / 480), or 214. Surveying 214 of the 480 employees would already have sufficed for ±5%; the census got 312 of them and, thanks to the correction, drops to ±3.2%.
Visualisations
A survey percentage is an interval. The "58% satisfied" of the example marks the centre of a range that the margin of error delimits: at 95% confidence, the true proportion of the population sits between 54.8% and 61.2%. The ±3.2% band placed around the percentage is the correct unit of reading, and reading the bare number amounts to mistaking the centre for the whole range. Two waves of the same survey that give 58% then 60% on samples of this size produce ranges that overlap widely, and it is that overlap which says whether the observed gap is a real movement or sampling noise.
Cost
| Phase | Level | Justification |
|---|---|---|
| Preparation | High | Everything is decided before the send: objective, definition of the population, sampling plan, drafting and proofreading the questions for neutrality, then the field test. A badly designed instrument is not corrected after the fact. |
| Execution | Low | Once the instrument is ready, distribution is nearly free and the marginal cost per additional respondent is close to zero. This is the technique's economic advantage over the interview. |
| Documentation | Medium | Closed questions are processed and tabulated cheaply; it is the open responses that weigh, coding free text being long and manual. The share of open questions decides this level. |
The cost profile is the inverse of the interview's: the survey is expensive to design and cheap to run, where the interview is quick to prepare but paid for at every repetition. This is what makes it the best use of the budget as soon as the population exceeds a few dozen people.
Tooling
The minimum tooling is a spreadsheet and an email: a questionnaire typed, sent as an attachment and processed by hand suits a small internal survey of a few dozen responses. It soon shows its limits, manual entry of the responses being both slow and error-prone.
Online forms (Microsoft Forms, Google Forms) are the everyday choice: they capture responses straight into a table, handle logic jumps and tabulate closed questions with no re-keying. Specialised survey platforms (Qualtrics, SurveyMonkey, Typeform, Tally) add what matters at greater scale: quota and panel management, distribution control, real-time response-rate tracking and targeted reminders. For an internal survey covering sensitive data, LimeSurvey deserves a mention of its own, because it is open source and can be hosted on Swiss infrastructure: the place where the responses are processed and stored is here a selection criterion at least as important as the features.
Downstream, the analysis splits by question type. Closed responses call for a spreadsheet for simple processing, then a statistical tool (R, Python, SPSS) as soon as margins of error must be computed, a significance tested or variables cross-tabulated. Open responses call for a qualitative coding tool (Dovetail, MAXQDA, NVivo, Atlas.ti) or a chain assisted by a language model for thematic grouping, provided the analyst reviews the categories produced. The thread that ties all of this together is traceability: every published proportion keeps the link to the question, the sample and the response rate it came from, without which the number outlives its own conditions of validity and starts to circulate on its own.
Sources
- IIBA, A Guide to the Business Analysis Body of Knowledge (BABOK Guide) v3, §10.45 Survey or Questionnaire: the definition, the two question types, the three elements (prepare, distribute, document) and the stated strengths and limits.
- IIBA, Guide to Business Data Analytics, §3.17 Survey or Questionnaire: the analytical angle, questions structured around a hypothesis, response-rate tracking for significance, anonymity and change feedback by functional unit.
- Cochran, W. G., Sampling Techniques, 3rd ed., John Wiley & Sons, 1977: the sample-size formula, the finite-population correction and the margin of error of a proportion.
- AAPOR, Standard Definitions: Final Dispositions of Case Codes and Outcome Rates for Surveys: the definition of the response rate on the eligible sample.
- AAPOR Task Force, Report on Non-Probability Sampling, 2013: the distinction between probability and non-probability sampling, then the one between response rate and participation rate.
- Pew Research Center, Writing Survey Questions: the effect of wording on the measurement, leading questions, scale order and the non-response of open questions.

