Premium Writing, Design & Advisory · Dubai, UAE

Labeeb Reference · every threshold traced to its primary source · last checked 22 September 2026

CFA and SEM Model Fit Indices: Thresholds, Sources and a Free Checker (2026)

The cut-offs quoted for CFI, RMSEA, SRMR, average variance extracted (AVE), composite reliability (CR) and HTMT come from a small number of papers. This page shows what each paper proposed, which values are conventions rather than rules, and a free checker that compares your AMOS or SmartPLS numbers with them.

Read this first: these are conventions, not laws

No fit index has an official pass mark. Every threshold below was proposed by a specific author, for specific conditions, and later authors disagree about some of them.

  • Hu and Bentler (1999) derived their cut-offs from simulation studies of particular model types. Marsh, Hau and Wen (2004) warned against treating them as “golden rules” for every model (source).
  • Henseler, Ringle and Sarstedt (2015) state that the exact HTMT threshold is debatable, which is why two values (.85 and .90) are in use (source).
  • Your supervisor, your university’s thesis guide or your target journal may set a different value. If they do, their rule is the one to follow, and the reason is worth citing in your methods chapter.
  • Covariance-based SEM (AMOS, lavaan, Mplus, LISREL) and PLS-SEM (SmartPLS) are assessed differently. Use the table that matches the method you ran.

If you are writing a thesis at a UAE or Saudi university, check your programme’s own research guide before your viva or submission. Where it names a threshold, cite the guide as well as the paper.

Covariance-based CFA and SEM: fit and validity criteria

For output from AMOS, lavaan, Mplus or LISREL. Each row names the paper that proposed the value. Full references with DOI links are at the end of the page.

Criterion Convention Primary source What to keep in mind
χ² (with df and p) p > .05: exact fit not rejected Kline (2023) Report χ², df and p even when p is small. With large samples the test rejects models with trivial misfit.
χ²/df (normed χ²) No agreed cut-off. About 5 was suggested early; stricter values of 2 or 3 appear in later textbooks Wheaton et al. (1977); Kline (2023) Kline cautions that the ratio has no clear-cut standard. Do not rely on it alone.
CFI ≥ .95 (close to .95); .90 is the older convention Hu & Bentler (1999); .90: Bentler & Bonett (1980) The .90 value was first proposed for Bentler and Bonett’s normed fit index and later carried over to CFI.
TLI (NNFI) ≥ .95 (close to .95) Hu & Bentler (1999) TLI is not bounded at 1. A value slightly above 1 is not an error.
RMSEA ≤ .06 (Hu & Bentler); ≤ .05 close and ≤ .08 reasonable (Browne & Cudeck); .08–.10 mediocre (MacCallum et al.) Hu & Bentler (1999); Browne & Cudeck (1992, 1993); MacCallum et al. (1996) Read it with its 90% confidence interval (next row).
RMSEA 90% CI and PCLOSE PCLOSE > .05: close fit (RMSEA ≤ .05) not rejected. An upper CI bound above .10 means poor fit cannot be ruled out Browne & Cudeck (1992, 1993); MacCallum et al. (1996) Report the interval, not just the point estimate. A wide interval often reflects a small sample.
SRMR ≤ .08 Hu & Bentler (1999) Hu and Bentler recommend reading SRMR together with CFI or RMSEA, not alone.
Standardised factor loadings ≥ .50, ideally ≥ .70 Hair et al. (2019), Multivariate Data Analysis A standardised loading above 1 points to an improper solution (Heywood case).
Composite reliability (CR) ≥ .60 (Bagozzi & Yi); ≥ .70 is the more common convention Bagozzi & Yi (1988); Hair et al. (2019) Bagozzi and Yi’s paper is the source usually cited for the .60 value.
AVE ≥ .50 Fornell & Larcker (1981); Bagozzi & Yi (1988) At .50 the construct explains at least as much item variance as measurement error does.
Fornell–Larcker criterion √AVE of each construct > its correlation with every other construct Fornell & Larcker (1981) Henseler et al. (2015) showed it often misses discriminant-validity problems. Report HTMT as well.
HTMT ≤ .85 strict; ≤ .90 for conceptually similar constructs; bootstrap CI should not include 1 Henseler, Ringle & Sarstedt (2015) Henseler et al. attribute .85 to Clark & Watson (1995) and Kline (2011), and .90 to Gold et al. (2001) and Teo et al. (2008).
Cronbach’s α ≥ .70 Nunnally (1978) α assumes equal loadings, which CFA rarely shows, so CR is usually reported alongside it. Taber (2018) questions treating .70 as a fixed rule.

PLS-SEM (SmartPLS): measurement and structural model criteria

PLS-SEM is assessed in two stages: the measurement model first, then the structural model. The values below come from Hair and colleagues’ open-access textbook (Hair et al., 2021), which follows their Primer on PLS-SEM, and from the papers they cite.

Criterion Convention Primary source What to keep in mind
Outer loadings ≥ .708; .40–.708 remove only if doing so raises CR or AVE; < .40 remove Hair et al. (2021), ch. 4 .708 is the value at which the construct explains 50% of the item’s variance.
Cronbach’s α, rho_A (ρA), composite reliability (ρC) .70–.90 satisfactory to good; above .90 problematic Hair et al. (2021), ch. 4; ρA: Dijkstra & Henseler (2015) α is the conservative bound and ρC the liberal one. ρA usually falls between them. Very high values suggest redundant items.
AVE ≥ .50 Fornell & Larcker (1981); Hair et al. (2021) Same logic as in covariance-based CFA.
HTMT ≤ .85 strict; ≤ .90 for conceptually similar constructs Henseler, Ringle & Sarstedt (2015) The recommended discriminant-validity test in PLS-SEM. Fornell–Larcker alone is not enough.
Inner VIF (collinearity) ≥ 5 probable collinearity; 3–5 possible; below 3 ideal Hair et al. (2021), ch. 6 Check this before you read path coefficients.
R² .75 substantial, .50 moderate, .25 weak Hair, Ringle & Sarstedt (2011); Hair et al. (2021), ch. 6 What counts as acceptable depends on the discipline. Some fields regularly report lower R².
f² (effect size) .02 small, .15 medium, .35 large Cohen (1988) Cohen proposed these values for multiple regression, and PLS-SEM applies them to each path.
Q² / Q²predict > 0 indicates predictive relevance Stone (1974); Geisser (1974); applied in Hair et al. (2019) SmartPLS reports Q²predict through PLSpredict. Hair et al. (2021) evaluate out-of-sample prediction with RMSE and MAE.

Free research calculator · runs in your browser

Average variance extracted (AVE) and composite reliability (CR) calculator

Average variance extracted (AVE) is the mean of the squared standardised loadings of a construct’s indicators: the share of item variance the construct explains rather than measurement error. It is the usual convergent-validity check in CFA and PLS-SEM (Fornell & Larcker, 1981), and the same AVE values feed the Fornell–Larcker discriminant-validity test.

For one reflective construct, paste its standardised indicator loadings from your CFA or PLS-SEM output, separated by commas or spaces. Use positive loadings after correctly reverse-coding any reverse-worded items. Do not use this calculation for formative indicators, unstandardised coefficients, cross-loadings or a model with unresolved improper solutions.

AVE = Σλ² / n. CR = (Σλ)² / [(Σλ)² + Σ(1 − λ²)], assuming standardised indicators with uncorrelated errors. Compare AVE with the CFA or PLS-SEM convention above; check your software output and reporting method before using a result. A value below .50 is a prompt to inspect the measurement model, not an instruction to delete items. See Fornell and Larcker (1981) and Hair et al. (2021).

If your AVE or CR needs interpretation alongside fit, discriminant validity or reviewer feedback, see specialist analysis support or discuss your model privately.

Free checker · runs in your browser, nothing is stored

Paste your numbers: CFA/SEM fit and validity checker

Pick the kind of model you ran, type the values from your AMOS, lavaan, Mplus, LISREL or SmartPLS output, and leave blank anything you do not have. Each value is compared with one published convention, and the source is shown next to it. The checker gives verdicts only. It does not write any text for your thesis.

Model type

Global fit

Measurement quality (enter the weakest value across your constructs)

Each verdict compares one number with one published convention. It is not a judgement of your model, and it is not a reason to add or drop items on its own. If your supervisor, university guide or target journal sets a different rule, follow that rule.

Fit below the thresholds, or AVE under 0.50? Send your model output from AMOS, SmartPLS or lavaan. A statistician reviews the measurement model and the defensible next steps, and explains each one so you can defend it at your viva.

Get your model reviewedData analysis support

What Hu and Bentler actually recommended, and dynamic fit

The familiar checklist (CFI ≥ .95, RMSEA ≤ .06 and SRMR ≤ .08, all passing together) is a simplification. Hu and Bentler (1999) tested two-index combination rules: a CFI or TLI close to .95 together with an SRMR close to .09, or an RMSEA close to .06 together with an SRMR close to .09. They also cautioned that no single cutoff works equally well across models and sample sizes.

Dynamic fit index cutoffs. McNeish and Wolf (2023) proposed simulating cutoffs for your own model (its number of items, its loadings and its sample size) instead of borrowing fixed values, and they provide a free web application to do it. Some journals and reviewers now ask for them. If your indices sit close to a fixed threshold, report the dynamic cutoffs alongside the conventional ones. Source: McNeish, D., & Wolf, M. G. (2023). Dynamic fit index cutoffs for confirmatory factor analysis models. Psychological Methods, 28(1), 61–88. PubMed record.

Where Labeeb comes in

A weak index is a question to investigate, not a verdict on your study. It can come from the model specification, a small sample, cross-loading items, a reverse-coded item that was never recoded, or simply the wrong cell copied from the output. Our analysts run and review SPSS, AMOS and SmartPLS analyses for students and researchers across the UAE, Saudi Arabia and the wider GCC. We check the data and the model, and we explain every table so you can defend it in your viva. We edit, analyse and explain. We do not write your thesis or your results chapter.

Related guides

Frequently asked questions

What are good CFI, TLI, RMSEA and SRMR values for a CFA?

The most cited values are CFI and TLI of .95 or above, RMSEA of .06 or below and SRMR of .08 or below (Hu & Bentler, 1999). Many studies still accept CFI and TLI from .90 and RMSEA up to .08, following Bentler & Bonett (1980) and Browne & Cudeck (1993). Report which convention you used and cite it.

Are Hu and Bentler’s cut-offs mandatory?

No. They came from simulations of particular models, and Marsh, Hau and Wen (2004) argued against applying them as universal rules. Treat them as widely used conventions. If your supervisor, university or journal specifies different values, follow them.

Should I use Fornell–Larcker or HTMT for discriminant validity?

Report both if your guide asks for Fornell–Larcker. Henseler, Ringle and Sarstedt (2015) showed that the Fornell–Larcker criterion often fails to detect a lack of discriminant validity, and proposed HTMT, with thresholds of .85 or .90.

Can Labeeb review my AMOS or SmartPLS model?

Yes. Our analysts can run or review the analysis, check the model and the data, and explain each result so you understand and can defend it. We do not write your thesis text: the interpretation you submit is yours.

How to cite this page (APA 7)

Labeeb Writing & Designs. (2026, September 22). CFA and SEM Model Fit Indices: Thresholds, Sources and a Free Checker (2026). https://labeeb.ae/cfa-sem-model-fit-thresholds/

Sources

Last checked: 22 September 2026. Each DOI below was checked on that date and resolves through doi.org. Books without a DOI are cited in full.

  1. Bagozzi, R. P., & Yi, Y. (1988). On the evaluation of structural equation models. Journal of the Academy of Marketing Science, 16(1), 74–94. https://doi.org/10.1007/BF02723327
  2. Bentler, P. M., & Bonett, D. G. (1980). Significance tests and goodness of fit in the analysis of covariance structures. Psychological Bulletin, 88(3), 588–606. https://doi.org/10.1037/0033-2909.88.3.588
  3. Browne, M. W., & Cudeck, R. (1992). Alternative ways of assessing model fit. Sociological Methods & Research, 21(2), 230–258. https://doi.org/10.1177/0049124192021002005. Reprinted in K. A. Bollen & J. S. Long (Eds.), Testing Structural Equation Models (1993, pp. 136–162). Sage.
  4. Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences (2nd ed.). Lawrence Erlbaum.
  5. Dijkstra, T. K., & Henseler, J. (2015). Consistent partial least squares path modeling. MIS Quarterly, 39(2), 297–316. https://doi.org/10.25300/MISQ/2015/39.2.02
  6. Fornell, C., & Larcker, D. F. (1981). Evaluating structural equation models with unobservable variables and measurement error. Journal of Marketing Research, 18(1), 39–50. https://doi.org/10.1177/002224378101800104
  7. Geisser, S. (1974). A predictive approach to the random effect model. Biometrika, 61(1), 101–107. https://doi.org/10.1093/biomet/61.1.101; and Stone, M. (1974). Cross-validatory choice and assessment of statistical predictions. Journal of the Royal Statistical Society: Series B, 36(2), 111–133. https://doi.org/10.1111/j.2517-6161.1974.tb00994.x
  8. Hair, J. F., Black, W. C., Babin, B. J., & Anderson, R. E. (2019). Multivariate Data Analysis (8th ed.). Cengage.
  9. Hair, J. F., Hult, G. T. M., Ringle, C. M., Sarstedt, M., Danks, N. P., & Ray, S. (2021). Partial Least Squares Structural Equation Modeling (PLS-SEM) Using R. Springer (open access). https://doi.org/10.1007/978-3-030-80519-7. Companion to Hair, Hult, Ringle & Sarstedt, A Primer on Partial Least Squares Structural Equation Modeling (PLS-SEM) (3rd ed., 2022). Sage.
  10. Hair, J. F., Ringle, C. M., & Sarstedt, M. (2011). PLS-SEM: Indeed a silver bullet. Journal of Marketing Theory and Practice, 19(2), 139–152. https://doi.org/10.2753/MTP1069-6679190202
  11. Hair, J. F., Risher, J. J., Sarstedt, M., & Ringle, C. M. (2019). When to use and how to report the results of PLS-SEM. European Business Review, 31(1), 2–24. https://doi.org/10.1108/EBR-11-2018-0203
  12. Henseler, J., Ringle, C. M., & Sarstedt, M. (2015). A new criterion for assessing discriminant validity in variance-based structural equation modeling. Journal of the Academy of Marketing Science, 43(1), 115–135. https://doi.org/10.1007/s11747-014-0403-8
  13. Hu, L., & Bentler, P. M. (1999). Cutoff criteria for fit indexes in covariance structure analysis: Conventional criteria versus new alternatives. Structural Equation Modeling, 6(1), 1–55. https://doi.org/10.1080/10705519909540118
  14. Kline, R. B. (2023). Principles and Practice of Structural Equation Modeling (5th ed.). Guilford Press.
  15. MacCallum, R. C., Browne, M. W., & Sugawara, H. M. (1996). Power analysis and determination of sample size for covariance structure modeling. Psychological Methods, 1(2), 130–149. https://doi.org/10.1037/1082-989X.1.2.130
  16. Marsh, H. W., Hau, K.-T., & Wen, Z. (2004). In search of golden rules: Comment on hypothesis-testing approaches to setting cutoff values for fit indexes and dangers in overgeneralizing Hu and Bentler’s (1999) findings. Structural Equation Modeling, 11(3), 320–341. https://doi.org/10.1207/s15328007sem1103_2
  17. Nunnally, J. C. (1978). Psychometric Theory (2nd ed.). McGraw-Hill. See also Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297–334. https://doi.org/10.1007/BF02310555
  18. Taber, K. S. (2018). The use of Cronbach’s alpha when developing and reporting research instruments in science education. Research in Science Education, 48, 1273–1296. https://doi.org/10.1007/s11165-016-9602-2
  19. Wheaton, B., Muthén, B., Alwin, D. F., & Summers, G. F. (1977). Assessing reliability and stability in panel models. Sociological Methodology, 8, 84–136. https://doi.org/10.2307/270754