Drone and eVTOL Noise Annoyance: From Laboratory to Community
Controlled listening tests show that some small-drone and synthetic electric vertical-takeoff-and-landing sounds produce more short-term annoyance than comparison sounds at equal A-weighted exposure or modeled loudness. Those results are increasingly cited as candidate source penalties, yet a 2026 NASA comparison found passenger-scale UAM vehicles and helicopters practically similar at equal A-weighted sound exposure across steady flight phases. Laboratory clips, contextual scenarios, and long-term community surveys also estimate different responses. We synthesized 53 screened sources, including 17 complete experiments, technical reports, reviews, and official surveys plus the complete public abstract of one current ISO standard. Field and demonstration evidence exists, but no screened source met all four criteria for a representative long-term operational community exposure-response study. Effects changed with vehicle class, comparator, reproduction, operation, ambient, event structure, and cueing. We therefore propose a four-class Laboratory-to-Community Translation Contract: rank-only, acute offset, contextual scenario, and community impact. Each upward claim requires explicit bridge validation for newly introduced calibration, source-path-receiver, event-pattern, contextual, sampling, and exposure dimensions. A listening-test offset is not a community standard until its bridges are observed.
Introduction
Drone and electric vertical-takeoff-and-landing vehicles introduce tonal, modulated, and often highly noticeable sounds into places that may have little conventional aircraft exposure. Controlled studies therefore ask a reasonable engineering question: when overall A-weighted energy or loudness is held constant, does a drone stimulus still produce a different short-term annoyance response? The answer matters for vehicle design, route planning, and certification, but it is not identical to the policy question of how many residents will be highly annoyed after months of real operations.
The evidence base has often allowed these two questions to blur. A calibrated NASA experiment found a systematic difference between the tested small uncrewed aircraft systems (sUAS) and road-vehicle recordings, yet its authors warned that the limited vehicles, event durations, and geometries could explain part or all of the offset. [[cite:christian2017]] In direct counterevidence, NASA's 2026 HULC test compared six passenger-scale UAM vehicles with four helicopters and found equal-ASEL annoyance offsets of about 1 dB or less across three steady flight phases. [[cite:hulc2026]] Later experiments associated annoyance with level or loudness and additional sound-quality descriptors, but their vehicle classes, playback systems, and metric implementations differed. [[cite:hui2021,metrics2022,boucher2024]] A laboratory result can therefore establish a source-specific acute effect without establishing a universal decibel penalty or a shared drone/eVTOL source class.
This paper asks: how far can short-term listening-test responses to drone and eVTOL sounds be transported into claims about real community multi-event exposure? It contributes a four-class Laboratory-to-Community Translation Contract . The contract separates within-domain ranking, acute equal-annoyance offsets, contextual scenario response, and population community impact. A claim may move upward only through an explicit bridge study that validates the newly introduced acoustic, temporal, contextual, and sampling dimensions. The contribution is not a new annoyance metric or a numerical penalty; it is a rule for keeping each output inside the evidence domain that produced it.
Methods
We performed a structured conceptual synthesis on 10 July 2026. Eight short English queries searched the AlexandrAI graph for direct duplicates and adjacent archive work. Twenty-nine external queries covered controlled drone listening tests, passenger-scale UAM comparisons, psychoacoustic models, auralization fidelity, visual and ambient context, repeated events, standards, community surveys, and the latest UK CAA update series. Fifty-three sources were screened from titles, abstracts, official pages, or full reports. Seventeen complete papers or reports were read in full. The ISO record was limited to, and cited as, its complete official public abstract; together these yielded 18 final references. Sources were included when they directly measured human response, specified response or reproduction validity, or defined the community estimand. Emission-only, hypothetical acceptance, inaccessible full-text, and conventional-aircraft sources were retained only as context unless they defined a necessary comparator.
For each full source we coded: acoustic input, reproduction and calibration, participant frame, response task, time horizon, event structure, context, allowed output, and the authors' stated limitations. We then followed citations forward from the foundational single-event test to its remote replication, from early soundscape work to later spatial-audio tests, from isolated-sound modeling to masking and multiple-event experiments, and from CAP 2505 through the 2024 and 2025 CAA updates. Contradictory evidence was sought deliberately: offsets that changed by stimulus subset or platform, sound-quality terms that changed sign or relevance, contextual cues with different effects, and event-number responses that departed from a single equal-energy rule.
Community-impact eligibility was audited with four explicit criteria: representative exposed residents (P), sustained mature real operations (O), measured or validated long-term home exposure (E), and standardized at-home annoyance (A). Every screened record carries at least one documented N exclusion certificate in the report data; unestablished dimensions are U and were never treated as failures. A C3 study would require P=O=E=A=Y. This first-failure design makes the corpus-bounded absence claim reproducible without pretending that every irrelevant source was fully coded on all four dimensions.
The synthesis treats an estimand as the precise response quantity a design can identify. A rating after a four-second clip estimates an acute response under that playback protocol. A four-minute sequence estimates scenario response over that event pattern. A standardized at-home survey paired with long-term exposure estimates population prevalence under sampled operations. Similar 0-10 response scales do not make these quantities interchangeable.
No quantitative meta-analysis was attempted. The experiments did not share vehicles, maneuvers, reference sounds, calibration, contexts, metric implementations, or participant sampling. Pooling their coefficients would create a precise average of non-equivalent quantities. Instead, evidence was synthesized by the largest claim class each design could support and by the bridge information missing for the next class.
Exposure Descriptors Are Not Response Outcomes
A-weighted equivalent level describes energy over a declared interval; sound exposure level compresses an event's A-weighted energy to a one-second reference. Loudness models perceived magnitude, while tonality, sharpness, roughness, and fluctuation strength represent different spectral or temporal sensations. None is itself a measured annoyance response. The distinction is explicit in the evidence: loudness and PNL often explained substantial variation, but additional associations depended on the tested vehicles and the model specification. [[cite:hui2021,metrics2022]]
A 578-participant pairwise-ranking study reinforces the conditionality. ISO 532-1 loudness best explained comparisons across 161 bench and 46 real-flight sounds, followed by A-weighted level; additional sharpness, tonality, roughness, or fluctuation associations changed with metric implementation and setup. The pairs were kept within 5.4 dB(A), not equalized in loudness, and were heard through participant-owned devices. [[cite:konig2024]] This is strong C0 ranking evidence for one drone family, not an isolated sound-quality penalty.
The TUSQ study makes the boundary concrete. Forty listeners rated 136 manipulated UAM sounds equalized to about 6 sones. The UAM set beat a shaped-noise reference in annoyance comparisons 73% of the time at equal loudness, yet only three of 141 total stimuli exceeded the sharpness threshold used by the candidate model and fluctuation strength occupied a narrow range. The fitted tonality-augmented model improved ranking within that dataset, but the authors explicitly called it far from a community-response predictor and noted that applicability to transient flyovers was unknown. [[cite:boucher2024]] Model output should therefore be labeled rank-screening within domain , not community annoyance.
Equal-Level Tests Establish Acute Source Effects
The strongest common result is modest but useful: level or loudness alone is not sufficient to order every tested drone sound. In the foundational NASA test, 38 participants heard 46 sUAS and 20 road-vehicle sounds in a calibrated three-dimensional room. The full-set augmented regression produced a median SELA offset of 5.64 dB, whereas a later subset reanalysis used for remote replication produced 4.14 dB. Both values described those samples, not a material constant. The original paper also observed that many sUAS events were longer and farther than the road events and stated that a duration-related correction could account for the significant offset. [[cite:christian2017,remote2023]]
Other designs support source-specific residual effects without agreeing on one mechanism. Hui and colleagues reproduced measured hover and flyby sounds from four quadcopters for 37 listeners and found strong level and loudness associations, while its distraction task reached a ceiling and was inconclusive. [[cite:hui2021]] Torija and Nicholls used 44 isolated four-second recordings from eight drone types; PNL and sharpness entered the selected annoyance model, but unknown participant hardware required pseudo-calibration and the authors analyzed relative response consistency rather than calibrated absolute exposure. [[cite:metrics2022]]
The variation is evidence against a universal adjustment, not evidence that the experiments failed. A source offset depends on the candidate metric, comparator, event window, vehicle, operation, reproduction chain, participant frame, and response question. Equal SEL removes one energy difference; equal loudness removes a modeled magnitude difference. Neither equalization makes duration, onset, tonality, modulation, motion, familiarity, or meaning equal.
Vehicle class is an especially important boundary. The larger offsets above came mainly from small quad- and multirotor stimuli or from a synthetic steady quadrotor family. HULC instead adopted a UAM definition of electric vertical-takeoff-and-landing aircraft with payloads between 800 and 8,000 lb, explicitly excluding sUAS. Its 40 participants rated 123 stimuli from six UAM vehicles and four helicopters. At equal ASEL, the UAM-helicopter offsets were 0.8 dB at departure, -0.4 dB at cruise, and -1.0 dB at approach; the authors judged the roughly 1 dB effects as not practically important. At equal observer distance, the tested UAM vehicles were 6 to 13 dB quieter in ASEL and about 2.4 to 2.6 annoyance points lower. [[cite:hulc2026]] Small-drone evidence cannot be relabeled as a passenger-eVTOL penalty.
The safe claim is therefore: several controlled stimulus sets show source-specific short-term annoyance variation not captured by A-weighted energy or loudness alone. The unsafe claim is: all drones require one fixed decibel penalty. A numerical offset is admissible only with its vehicle and comparator set, integration window, playback calibration, participant sample, uncertainty interval, and response task attached.
Playback and Context Change the Estimand
Playback validity is not a file-format property. A listening stimulus represents a source only if the source spectrum and directivity, propagation, motion, receiver, reproduction system, and level calibration are adequate for the intended claim. It represents a soundscape only if the ambient environment and contextual cues are also plausible. NASA guidance therefore calls for matched recording and reproduction methods, spatial metadata, calibrated levels, and ambient recordings that represent both indoor and outdoor target communities. [[cite:begault2021]]
The remote NASA replication shows why platform is part of the estimand. Its 48 completers broadly reproduced the ordering of the most and least annoying sounds and the qualitative sUAS-road separation, but participant-owned hardware produced a larger offset than the calibrated room. The hand-rub calibration pilot itself had a 4.3 dBA standard deviation. A cue asking listeners to imagine repeated sounds outdoors near home also raised sUAS ratings. [[cite:remote2023]] Remote tests may support ranking and efficient screening; they do not inherit absolute exposure control merely because the audio is lossless.
Ambient context can change both incremental and absolute response. In a calibrated audio-visual experiment, the same 65 dBA hovering quadcopter increased annoyance far more in low-traffic scenes than in road-traffic-heavy scenes, and adding panoramic visuals changed annoyance and pleasantness more clearly than loudness. [[cite:torija2020]] A later spatial-audio experiment likewise found larger incremental UAS effects in a quieter environment, but the absolute annoyance of the complete quiet scene could still be below that of the busy scene. [[cite:lotinga2025]] Thus, 'larger incremental effect in quiet' is not equivalent to 'higher absolute annoyance in quiet.'
Meaning is also conditional. A 2025 online experiment found lower ratings when participants received extensive medical-delivery context, but it could not relate those ratings to absolute playback level. Its authors noted that another recent study found no significant medical-context effect. [[cite:woodcock2025]] Context is neither noise nor nuisance to be averaged away; it is a component of the response condition that requires replication before it becomes an operational mitigation claim.
Multiple Events Introduce a New Response Function
A single-event rating does not specify how annoyance grows when events repeat. Equal-energy accounting assigns a 10 dB increase when the number of equal exposures grows tenfold, but psychological response may depend on whether events remain individually distinct, blend into a continuous stream, interrupt activity, or alter quiet intervals. Event number is therefore not just another column in a single-event regression.
NASA's Noise and Number test exposed 38 participants to nine four-minute spatialized UAM scenarios varying gain and rate. A single tradeoff near 13 dB per tenfold event increase fit the means better than the equal-energy value of 10 dB; a step model around two operations per minute fit more closely, consistent with a transition from discrete events toward a blended sound. The authors described the response growth as potentially non-constant and dependent on the sound and context. [[cite:multievent2023]] This is direct evidence that event structure can matter, but it is an acute four-minute result, not an annual dose-response curve.
Masking adds another conditional transition. Laboratory data support a discount when a target approaches partial masking, whereas clearly audible targets receive little discount. The fitted relationship depends on target-to-masker relation and listener response rather than a fixed urban-versus-rural subtraction. [[cite:tracy2024]] Operational models must therefore preserve event timing, quiet gaps, ambient evolution, and noticeability instead of collapsing an entire schedule to one level before response validity has been established.
Community Annoyance Is a Population Estimand
The public abstract of ISO/TS 15666 concerns social surveys of annoyance at home . It states that the specification covers questions, response scales, key conduct, and reporting to improve comparability, while warning that conformity does not guarantee accurate prevalence or an interpretable exposure relationship; sampling and exposure uncertainty remain decisive. [[cite:iso15666]] A clip rating can borrow an 11-point scale without becoming the same outcome.
The FAA Neighborhood Environmental Survey illustrates the additional design layers. More than 10,000 residents near 20 representative airports answered a disguised multi-topic mailed survey about their previous 12 months at home. Modeled 2015 conventional-aircraft DNL was paired with the responses, and the resulting national curve showed higher annoyance than the legacy Schultz curve. [[cite:faa_nes]] This is evidence that community response changes with fleet, period, place, survey, and population. It is not evidence that the conventional-aircraft curve transfers to low-altitude drones.
Field and operational-adjacent evidence does exist. CAP 2962 reports a limited pre/post perception survey around Isleworth drone flights, with 239 first-phase responses but no linked acoustic dose or standardized noise-annoyance outcome. CAP 3086 reports a 16-person live-drone soundwalk and 52 responses from 36 citizens during three-day flight demonstrations; the respondents were not a representative exposed-resident cohort. [[cite:caa2962,caa3086]] These studies provide bridge evidence, not a population community exposure-response function.
No screened source, including the UK CAA annual updates through March 2025, met all four C3 criteria: representative residents exposed to sustained routine drone or eVTOL operations, measured or validated long-term home exposure, and standardized at-home annoyance. [[cite:torija_clark2021,caa2962,caa3086]] Experimental single-event exposure-response curves exist; the missing object is a long-term operational community curve. The corpus-bounded absence does not imply zero impact or justify a precautionary penalty of arbitrary size. It means that current evidence identifies design risks and testable mechanisms while population calibration remains unobserved.
The Laboratory-to-Community Translation Contract
The proposed contract classifies every result by the largest claim it can identify. The four classes are cumulative but not automatically nested: each upward move introduces a new bridge that must be tested. A higher class may reuse lower-class acoustic evidence, but it cannot inherit validity by terminology, a shared response scale, or a more elaborate metric.
Admissible(C k ) = 1 only if E 0..k and B 0->1..k-1->k are documented
The bridges are falsifiable. The C0-to-C1 bridge fails if relative rankings change under calibrated playback or if level-response overlap is insufficient. The C1-to-C2 bridge fails if the offset changes with ambient, event duration, movement, or cueing beyond its declared uncertainty. The C2-to-C3 bridge fails if scenario predictions do not calibrate against exposed residents or if selection and exposure error dominate the relationship. A failed bridge narrows the claim; it does not erase the lower-class result.
Every published coefficient should therefore carry a claim-class label and a domain card: vehicle and operation, acoustic descriptor and implementation, event window, calibration and reproduction, ambient and visual condition, participant frame, response horizon, comparator, uncertainty, and known bridge failures. This turns 'drone annoyance' from an ambiguous endpoint into an auditable chain of estimands.
Discussion
The research question has a bounded answer. Controlled tests can support relative ranking and, when calibrated with common comparators, stimulus-specific acute offsets. Contextual and multiple-event experiments can support response surfaces for the scenarios they reproduce. None of these designs alone estimates the prevalence of long-term at-home annoyance under real operations. The translation contract makes that boundary explicit without dismissing laboratory evidence.
For vehicle designers, C0 and C1 remain valuable. Equal-loudness comparisons can reveal tonal or modulation variants worth redesigning, and repeated calibrated tests can identify robust within-family improvements. The NASA TUSQ work is appropriately positioned as short-term sound-design modeling, not regulation. [[cite:boucher2024]] HULC additionally shows that ASEL, loudness exposure, perceived-noise, and psychoacoustic-annoyance metrics achieved very similar within-study fits; no metric gained a decisive class-level advantage. [[cite:hulc2026]] For route and operations designers, C2 is the relevant target: altitude, operation, event rate, quiet intervals, ambient masking, and service cues must be assembled into scenarios whose acoustic rendering is validated.
For regulators, an energy descriptor can remain necessary for exposure accounting without being sufficient for response. Adding an unvalidated fixed penalty would hide rather than solve transportability. A better interim record reports energy metrics alongside event counts, timing, maximum levels, source-specific sound-quality descriptors, and the claim class supported by human-response evidence. The purpose is not to replace one scalar with many decorative metrics; it is to retain the dimensions needed for later field calibration.
For community studies, staged deployment creates a validation opportunity. Pre-operation baseline soundscapes, transparent route schedules, indoor and outdoor monitoring, repeated at-home surveys, and nonresponse follow-up can connect C2 predictions to C3 outcomes. Designs should pre-register whether they estimate absolute annoyance, change from baseline, highly-annoyed prevalence, sleep disturbance, or acceptance. Those outcomes may be related, but none should be substituted for another.
Limitations and Validation Agenda
This synthesis has six limitations. First, it did not estimate a pooled effect because the source experiments were intentionally heterogeneous and because sUAS, cargo drones, and passenger-scale UAM vehicles are not one exchangeable source class. Second, several recent paywalled studies were screened but excluded from load-bearing claims when full text could not be verified. Third, the field changes quickly; the search date bounds the corpus. Fourth, the proposed contract is a conceptual method and has not yet been prospectively evaluated against a mature operational community dataset. Fifth, community annoyance is only one impact; sleep, cognition, health, equity, wildlife, and acceptance require separate outcomes and evidence chains. Sixth, CAP 2962 and CAP 3086 are narrative annual updates rather than systematic reviews, and the latest dedicated CAA human-effects window found ends in March 2025.
A decisive validation program would use at least three linked protocols. Protocol A would reproduce a vehicle family's rank and acute offset across two calibrated laboratories and a controlled remote platform. Protocol B would cross the same sounds with measured urban and rural ambients, indoor facade transfer, visual scenes, and event schedules, then test held-out scenario prediction. Protocol C would follow residents before and after a staged route begins, combine validated exposure reconstruction with ISO-compatible at-home outcomes, quantify nonresponse and route-selection bias, and compare observed responses with Protocol B forecasts.
The primary validation statistic should be bridge calibration, not only within-study R-squared. For C1-to-C2, report prediction error and calibration slope across held-out contexts and event patterns. For C2-to-C3, report calibration of absolute prevalence and change from baseline across neighborhoods, with uncertainty from exposure and sampling. A metric that ranks clips well but miscalibrates neighborhoods remains a useful C0 tool and a failed C3 model.
Conclusion
Drone-noise listening tests have established a real but limited result: several tested small-drone and synthetic UAM sound sets produce short-term annoyance differences that A-weighted energy or loudness alone does not fully capture. Passenger-scale evidence does not reproduce a large class penalty: HULC found UAM and helicopter responses nearly identical at equal ASEL across steady departure, cruise, and approach. [[cite:hulc2026]] The size and apparent mechanism vary with source class, stimulus construction, vehicle and maneuver, comparator, reproduction, ambient context, event pattern, cueing, and model. That variation rules out a universal penalty while preserving the value of controlled sound-design evidence.
The Laboratory-to-Community Translation Contract keeps the inference proportional to the design. Rank-only evidence supports screening; calibrated comparisons support protocol-specific acute offsets; validated contextual sequences support scenario response; representative long-term at-home studies support community exposure-response. Each transition requires its own bridge. Until those bridges are observed, a listening-test coefficient should remain a laboratory result, not a community standard.