# Do your evaluations have enough power?

**URL:** <https://community.epinowcast.org/t/do-your-evaluations-have-enough-power/366>\
**Category:** Uncategorized\
**Created:** [15 December 2025 17:03 UTC](https://community.epinowcast.org/t/do-your-evaluations-have-enough-power/366 "2025-12-15T17:03:46Z")\
**Posts on this page:** 1\
**Showing post:** 8

<div class="post-metadata">

**Author:** ![sbfnk](https://dub1.discourse-cdn.com/flex005/user_avatar/community.epinowcast.org/sbfnk/32/28_2.png) [@sbfnk](https://community.epinowcast.org/u/sbfnk)\
**Post date:** [13 January 2026 15:54 UTC](https://community.epinowcast.org/t/do-your-evaluations-have-enough-power/366/8 "2026-01-13T15:54:49Z")

</div>

This is a great question - my sense from all the hub etc. work is that it’s likely underpowered but it would be great to think about this a bit more, and even more so to develop some guidance on how to address this when reporting forecast scores.

One thing that I think we discussed in the past was the idea of [Model Confidence Sets](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=522382), i.e. sets of models that are indistinguishable in their forecast ability - there seems to be some active work on this, with applications to COVID forecasts in [Sequential model confidence sets](https://arxiv.org/abs/2404.18678v3) and to forecasts during particular phases in [Conditional model confidence sets](https://arxiv.org/abs/2505.21278v1).

---

_[View the full topic](https://community.epinowcast.org/t/do-your-evaluations-have-enough-power/366)._
