All guides

Analysis and reporting

Reporting Likert data in a manuscript: when means are fine and when they are not

Supratik Panuganti
Co-founder & CTO, Sutrix
September 18, 2026 · 6 min read

Short answer

Treat a single Likert item as ordinal: report counts and percentages per response option, use the median for a summary, and use rank-based tests for comparisons. Treat a multi-item Likert scale, meaning a sum or mean across several items, as approximately continuous: means, standard deviations, and parametric tests are acceptable and widely used. The reviewer objection you want to avoid is averaging a single five-point item and calling the result a 3.7.

An item is not a scale

Rensis Likert's original method summed responses across many items to produce a scale score. A single item with five ordered options is a Likert-type item, and it is ordinal: the distance between "agree" and "strongly agree" is not known to equal the distance between "neutral" and "agree". A scale built from several such items behaves much more like a continuous measurement, and decades of simulation work show that parametric methods perform well on it.

Most disputes in review come from blurring this line. Decide, item by item, whether you are reporting an individual question or a validated multi-item score, and let that decision drive the statistics.

Reporting a single item

Show the distribution. A table or a stacked bar chart with the count and percentage for each response option tells the reader more than any summary statistic. If a single summary is needed, use the median and interquartile range. Comparing two groups on one item calls for a rank-based test such as the Mann-Whitney U test, or an ordinal regression when you need to adjust for covariates.

Collapsing responses, for example agree plus strongly agree versus everything else, is acceptable when the cut is defined in advance and stated. Choosing the cut after looking at the data is not.

Reporting a multi-item scale

For a validated scale score, report the mean and standard deviation, the number of items and the possible range, and the internal consistency in your sample, usually Cronbach's alpha. Group comparisons with t-tests or analysis of variance, and adjusted analyses with linear regression, are standard. Check the distribution for ceiling or floor effects, which are common in symptom scales, and mention them if present.

If the scale has published norms or T-score conversions, such as PROMIS measures, report the converted score and cite the conversion source so readers can compare across studies.

Presentation choices that prevent review comments

State in the methods how each variable was treated and why, in one sentence. Report exact response wording for any item you analyze individually. Give effect sizes with confidence intervals rather than p-values alone. And keep the decimal places honest: a median response on a five-point item is a whole number or a half, not 3.72.

  • Single item: counts and percentages per option, median and IQR, rank-based tests or ordinal regression.
  • Multi-item scale: mean and SD, range, internal consistency, parametric tests or linear models.
  • Pre-specify any collapsing of categories.
  • Report effect sizes with confidence intervals.

Common questions

Is it ever acceptable to report a mean for a single Likert item?

Some fields do it, and with large samples the practical harm is small, but it invites a reviewer comment. If you report it, also report the distribution, and be prepared to justify the choice.

How many items make a scale continuous enough for parametric tests?

There is no strict cutoff. Simulation studies show parametric tests are robust with as few as four or five items summed, provided the distribution is not badly skewed. Validated instruments have already settled this question in their development papers.

Should I use non-parametric tests to be safe?

Not by default. Rank-based tests on a well-behaved scale score lose a little power and make effect sizes harder to interpret. Match the method to the measurement level instead.

Sources

  1. 1.Sullivan GM, Artino AR. Analyzing and interpreting data from Likert-type scales. J Grad Med Educ. 2013.
  2. 2.Norman G. Likert scales, levels of measurement and the laws of statistics. Adv Health Sci Educ. 2010.
  3. 3.Carifio J, Perla R. Resolving the 50-year debate around using and misusing Likert scales. Med Educ. 2008.

Let’s talk about your study

What do you want
your data to answer?

Bring your research question and where the analysis is getting stuck. We’ll discuss the data you have, the work you need, and whether Sutrix is a fit.

Book a study call

30 minutes with Supratik, co-founder · Google Meet

How is my project priced?

We scope the work around your data, instruments, research questions, and the materials you need. We agree the deliverables, timing, project price, and included follow-up before analysis begins. The call is where we establish what your study needs.

Who checks the analysis?

Sutrix is a managed service with human review. Our team checks scoring, analytical choices, and interpretation; your team supplies the study context. The delivered methods and code let you inspect the work. Meet the founders.

What should I bring to the call?

Your research question, the kind of data you collected, where you’re stuck, and any deadline. No account or dataset is needed. We agree on data-sharing arrangements before requesting de-identified study data. Read about data handling.