# Streamlining of epi modeling tools

**URL:** <https://community.epinowcast.org/t/streamlining-of-epi-modeling-tools/241>\
**Category:** Uncategorized\
**Created:** [14 August 2024 15:31 UTC](https://community.epinowcast.org/t/streamlining-of-epi-modeling-tools/241 "2024-08-14T15:31:18Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![adrianlison](https://dub1.discourse-cdn.com/flex005/user_avatar/community.epinowcast.org/adrianlison/32/27_2.png) [@adrianlison](https://community.epinowcast.org/u/adrianlison)\
**Post date:** [14 August 2024 15:31 UTC](https://community.epinowcast.org/t/streamlining-of-epi-modeling-tools/241/1 "2024-08-14T15:31:18Z")

</div>

The following posts are spun out of the discussion from [Community Seminar 2024-08-07 - Kaitlyn Johnson - Wastewater modeling to forecast hospital admissions in the US: Challenges and opportunities](https://community.epinowcast.org/t/community-seminar-2024-08-07-kaitlyn-johnson-wastewater-modeling-to-forecast-hospital-admissions-in-the-us-challenges-and-opportunities/237) as a separate discussion thread (see below).

---

<div class="post-metadata">

**Author:** ![adrianlison](https://dub1.discourse-cdn.com/flex005/user_avatar/community.epinowcast.org/adrianlison/32/27_2.png) [@adrianlison](https://community.epinowcast.org/u/adrianlison)\
**Post date:** [14 August 2024 15:31 UTC](https://community.epinowcast.org/t/streamlining-of-epi-modeling-tools/241/2 "2024-08-14T15:31:48Z")

</div>

@jamesazam wrote:

> [@Community Seminar 2024-08-07 - Kaitlyn Johnson - Wastewater modeling to forecast hospital admissions in the US: Challenges and opportunities](https://community.epinowcast.org/t/community-seminar-2024-08-07-kaitlyn-johnson-wastewater-modeling-to-forecast-hospital-admissions-in-the-us-challenges-and-opportunities/237/6):
>
> Thanks for your fantastic talk yesterday @kejohnson9. I was particularly interested in one of your leading questions, i.e., “How do we streamline deploying scientific ideas/models into production and tool development”? I think the forecasting/nowcasting community is fast evolving and it’s important to act on this now rather than later. On the call, I asked if we could employ the [tidymodels](https://www.tidymodels.org/) approach where the community has [developed guidelines for developing models](https://tidymodels.github.io/model-implementation-principles/). Some interface/engine/output …

---

<div class="post-metadata">

**Author:** ![adrianlison](https://dub1.discourse-cdn.com/flex005/user_avatar/community.epinowcast.org/adrianlison/32/27_2.png) [@adrianlison](https://community.epinowcast.org/u/adrianlison)\
**Post date:** [14 August 2024 15:32 UTC](https://community.epinowcast.org/t/streamlining-of-epi-modeling-tools/241/3 "2024-08-14T15:32:09Z")

</div>

@kejohnson9 replied:

> [@Community Seminar 2024-08-07 - Kaitlyn Johnson - Wastewater modeling to forecast hospital admissions in the US: Challenges and opportunities](https://community.epinowcast.org/t/community-seminar-2024-08-07-kaitlyn-johnson-wastewater-modeling-to-forecast-hospital-admissions-in-the-us-challenges-and-opportunities/237/8):
>
> On the call, I asked if we could employ the [tidymodels](https://www.tidymodels.org/) approach where the community has [developed guidelines for developing models](https://tidymodels.github.io/model-implementation-principles/). I think this is a really great point, and I completely agree with your assessment that taking the time now to streamline these would have a huge impact going forward. As someone relatively new to this, I would say it would also be really helpful because it would remove some of these decisions about interface/formatting inputs and outputs that we face as developer…

---

<div class="post-metadata">

**Author:** ![adrianlison](https://dub1.discourse-cdn.com/flex005/user_avatar/community.epinowcast.org/adrianlison/32/27_2.png) [@adrianlison](https://community.epinowcast.org/u/adrianlison)\
**Post date:** [14 August 2024 15:32 UTC](https://community.epinowcast.org/t/streamlining-of-epi-modeling-tools/241/4 "2024-08-14T15:32:33Z")

</div>

@jamesazam replied:

> [@Community Seminar 2024-08-07 - Kaitlyn Johnson - Wastewater modeling to forecast hospital admissions in the US: Challenges and opportunities](https://community.epinowcast.org/t/community-seminar-2024-08-07-kaitlyn-johnson-wastewater-modeling-to-forecast-hospital-admissions-in-the-us-challenges-and-opportunities/237/9):
>
> Yes, exactly. For example, there is a similar conversation happening in EpiNow2 about restructuring the [output class](https://github.com/epiforecasts/EpiNow2/issues/451) and returning the summarised output in a [ready format for use with scoringutils](https://github.com/epiforecasts/EpiNow2/issues/618). Here, scoringutils could dictate the structure of outputs but we must be wary of tightly coupling these tools instead of making them interoperable. See a similar restructuring exercise in [serofoi](https://github.com/epiverse-trace/serofoi/issues/89). Other standardisations could include function and input naming conventions, etc.

---

<div class="post-metadata">

**Author:** ![adrianlison](https://dub1.discourse-cdn.com/flex005/user_avatar/community.epinowcast.org/adrianlison/32/27_2.png) [@adrianlison](https://community.epinowcast.org/u/adrianlison)\
**Post date:** [14 August 2024 15:32 UTC](https://community.epinowcast.org/t/streamlining-of-epi-modeling-tools/241/5 "2024-08-14T15:32:53Z")

</div>

@kejohnson9 replied:

> [@Community Seminar 2024-08-07 - Kaitlyn Johnson - Wastewater modeling to forecast hospital admissions in the US: Challenges and opportunities](https://community.epinowcast.org/t/community-seminar-2024-08-07-kaitlyn-johnson-wastewater-modeling-to-forecast-hospital-admissions-in-the-us-challenges-and-opportunities/237/10):
>
> Sam had sent me this issue a few days ago, and this is exactly what we’re trying to decide now. See issue [here](https://github.com/CDCgov/ww-inference-model/issues/49). Have you all landed on a set of outputs coming from a wrapper fitting function? I really liked the idea of passing back the stan\_args, but also think we need to pass back the input data with the correct metadata mapped to it (e.g. the dates for example). I would also be interested in making the summarized output in a format readily usable for scoringutils, we haven’t quite gotten to …

---

<div class="post-metadata">

**Author:** ![adrianlison](https://dub1.discourse-cdn.com/flex005/user_avatar/community.epinowcast.org/adrianlison/32/27_2.png) [@adrianlison](https://community.epinowcast.org/u/adrianlison)\
**Post date:** [14 August 2024 15:33 UTC](https://community.epinowcast.org/t/streamlining-of-epi-modeling-tools/241/6 "2024-08-14T15:33:40Z")

</div>

[jamesazam](https://community.epinowcast.org/u/jamesazam) replied:

> Have you all landed on a set of outputs coming from a wrapper fitting function? I really liked the idea of passing back the `stan_args`, but also think we need to pass back the input data with the correct metadata mapped to it (e.g. the dates for example).

I’m not sure we’ve agreed yet 😅.

> I would also be interested in making the summarized output in a format readily usable for scoringutils, we haven’t quite gotten to any evaluation modules in the package yet but intend to.

I think that would be quite convenient as users wouldn’t have to wrangle further. Your package could have a custom as.forecast method as suggested by Sam.

---

<div class="post-metadata">

**Author:** ![adrianlison](https://dub1.discourse-cdn.com/flex005/user_avatar/community.epinowcast.org/adrianlison/32/27_2.png) [@adrianlison](https://community.epinowcast.org/u/adrianlison)\
**Post date:** [14 August 2024 15:34 UTC](https://community.epinowcast.org/t/streamlining-of-epi-modeling-tools/241/7 "2024-08-14T15:34:21Z")

</div>

@samabbott replied:

> [@Community Seminar 2024-08-07 - Kaitlyn Johnson - Wastewater modeling to forecast hospital admissions in the US: Challenges and opportunities](https://community.epinowcast.org/t/community-seminar-2024-08-07-kaitlyn-johnson-wastewater-modeling-to-forecast-hospital-admissions-in-the-us-challenges-and-opportunities/237/12):
>
> Here is the EpiNow2 issue ([Return forecasts for easy processing with `scoringutils` · Issue #618 · epiforecasts/EpiNow2 · GitHub](https://github.com/epiforecasts/EpiNow2/issues/618)) and here ([Update scoringutils integration to offer a as\_forecast\_samples.epinowcast method · Issue #455 · epinowcast/epinowcast · GitHub](https://github.com/epinowcast/epinowcast/issues/455)) is the epinowcast issue. In the new scoringutils 2.0.0 we would just need a as\_forecast\_sample method and then you can map to all the other data formats you might want (i.e quantiles - I know nice work @nikosbosse). oops looks l…

---

<div class="post-metadata">

**Author:** ![adrianlison](https://dub1.discourse-cdn.com/flex005/user_avatar/community.epinowcast.org/adrianlison/32/27_2.png) [@adrianlison](https://community.epinowcast.org/u/adrianlison)\
**Post date:** [14 August 2024 15:36 UTC](https://community.epinowcast.org/t/streamlining-of-epi-modeling-tools/241/8 "2024-08-14T15:36:29Z")

</div>

@samabbott replied:

> I think some of the really obvious wins for sharing tools are:
> 
> - Specifying priors and making it easy to pass to stan in some generalised way. I think `brms` functionality could be adapted here.
> - Sharing a formula interface. My hope here had been again to leverage `brms` or similar but I couldn’t make it work. `epinowcast` currently has a custom interface and either spinning that out or make a new version that is easy to pass to stan would really help.
> - Delay distribution handling from estimating through to handling (for example correct discretisation). @athowes has done a lot of work on the flexible estimation side (extending `brms`) and again `epinowcast` has some custom stuff for discretising but again that could be improved and generalised.
> - Better integration with `tidybayes` would also I think help a lot for a lot of these tools (again @athowes is exploring this in `epidist`).
> - I think it would be quite easy to create a shared library of stan functions where people depend on them using git submodules. Its not ideal but I think it would work. A more complex version would transpile this out to C++ and then this could be more easily distributed and integrated into stan.
> - Pre and post processing data. Again in `epinowcast` we found a lot of gotchas in processing real time data that I think are often missed.
> - Visualisation. With the new `forecast_samples` class in `scoringutils` I think a really nice model plotting package could be made. Potentially an Rt specific version could also be made.
> - @pearsonca has a few projects like `coerceDT` that aim to make some of the things we commonly do (here verify data.table inputs) easier and more robust.
> 
> There are also lots of others that could be done without needing to learn new things/ massively engineer.
> 
> I am currently working on (with @sambrand) a new Julia ecosystem ([EpiAware.jl: Real-time infectious disease monitoring · EpiAware.jl](https://cdcgov.github.io/Rt-without-renewal/dev/) - excuse the LLM text in some of the docs) where I think making these interactive modules should be a lot easier. That is still in the design stages though and obviously Julia doesn’t have the uptake or dependability that R and stan has.

---

<div class="post-metadata">

**Author:** ![adrianlison](https://dub1.discourse-cdn.com/flex005/user_avatar/community.epinowcast.org/adrianlison/32/27_2.png) [@adrianlison](https://community.epinowcast.org/u/adrianlison)\
**Post date:** [14 August 2024 15:36 UTC](https://community.epinowcast.org/t/streamlining-of-epi-modeling-tools/241/9 "2024-08-14T15:36:54Z")

</div>

@kejohnson9 replied:

> [@Community Seminar 2024-08-07 - Kaitlyn Johnson - Wastewater modeling to forecast hospital admissions in the US: Challenges and opportunities](https://community.epinowcast.org/t/community-seminar-2024-08-07-kaitlyn-johnson-wastewater-modeling-to-forecast-hospital-admissions-in-the-us-challenges-and-opportunities/237/14):
>
> @seabbs I am curious which of these seem the highest priority/lowest lift and could represent a proof of concept/ test of the added value (or if you think thats silly and we have to do all of them at once). For example, my instinct is that a shared interface for the inputs (so prior specification, delay distribution handling, and preprocessing) would be extremely helpful for evaluation/identifying which models work best in which contexts all else remaining equal. @jamesazam will be following t…

---

<div class="post-metadata">

**Author:** ![adrianlison](https://dub1.discourse-cdn.com/flex005/user_avatar/community.epinowcast.org/adrianlison/32/27_2.png) [@adrianlison](https://community.epinowcast.org/u/adrianlison)\
**Post date:** [14 August 2024 15:37 UTC](https://community.epinowcast.org/t/streamlining-of-epi-modeling-tools/241/10 "2024-08-14T15:37:19Z")

</div>

@samabbott replied:

> Definitely don’t need to do them all at once (and I think by its nature the modular approach is the way to go).
> 
> I think the formula interface that `epinowcast` uses could be spun out and I imagine that could then be used (after maybe a refactor to improve) in say `EpiSewer` fairly easily (@adrianlison ?). The only blocker to that is whether or not it makes sense to spin it out or if something like `brms` actually could be used (me not managing it doesn’t mean it isn’t very doable).
> 
> I also think the share stan library is very doable with some effort and thought as long as people are willing to take the submodule dependency approach.
> 
> The work on distribution specification that @sbfnk and @jamesazam have been doing also seems like it could fairly easily be spun out and usefully used elsewhere in the very near term.
> 
> I think all these things are mostly about resource and will as none of them are on the shortest path to any outcome and it has proved extremely difficult to convince funders etc they are important.

---

<div class="post-metadata">

**Author:** ![adrianlison](https://dub1.discourse-cdn.com/flex005/user_avatar/community.epinowcast.org/adrianlison/32/27_2.png) [@adrianlison](https://community.epinowcast.org/u/adrianlison)\
**Post date:** [14 August 2024 15:37 UTC](https://community.epinowcast.org/t/streamlining-of-epi-modeling-tools/241/11 "2024-08-14T15:37:54Z")

</div>

@adrianlison replied:

> What would be a real game changer for me is to have a domain (but not inference tool)-specific language for the kinds of semi-mechanistic generative epi models we use + a package to represent this in a convenient data structure **in `R`.** This is because I expect `R` to remain popular in epi research and public health for quite some time, while probabilistic programming is evolving fast and I don’t know how long people will still like `stan`.
> 
> If there was a ppl-agnostic structure to represent the different models in `epinowcast`, `EpiNow2`, `EpiSewer`, `epidist`, etc. (and I think that @samabbott and @sambrand have already laid the conceptual groundwork for this in `EpiAware`, although in `Julia` and maybe still a bit tightly coupled with `Turing`), and a package with helper functions to produce corresponding `stanargs` (including inits and priors, as mentioned above) that follow some clear convention, then we would as a first step “only” have to
> 
> - adjust our `stan` models to receive arguments for the “data” block from this representation
> - adjust the interface functions of our `R` packages to define the inputs for `stan` via this ppl-agnostic representation (instead of directly specifying stan arguments as we now do most of the time)
> 
> This would not necessarily require a shared stan library as it leaves the ppl implementation of the models completely open (it remains the responsibility of the developer to implement the model correctly and to reject representations that your stan model cannot implement).
> 
> The immediate advantage of this would be cleaner interface functions in `R` with a lower entry barrier for new contributors - right now, I can at least say for `epinowcast` and `EpiSewer` that IMO you need quite an in-depth understanding of the respective stan model (including variable names!) to work on the interface functions… which is really bad.
> 
> The more long-term advantage of this would be that this could pave the way for transitioning to other ppls (or offering several backends) _while still offering our tools in R_. I guess that in `Julia` you could ideally take such a representation and directly construct the respective EpiAware model during runtime. And if someone wants to try build that with `stan`/`brms`, we won’t stop them 😃

---

<div class="post-metadata">

**Author:** ![adrianlison](https://dub1.discourse-cdn.com/flex005/user_avatar/community.epinowcast.org/adrianlison/32/27_2.png) [@adrianlison](https://community.epinowcast.org/u/adrianlison)\
**Post date:** [14 August 2024 15:38 UTC](https://community.epinowcast.org/t/streamlining-of-epi-modeling-tools/241/12 "2024-08-14T15:38:28Z")

</div>

@samabbott replied:

> [@Community Seminar 2024-08-07 - Kaitlyn Johnson - Wastewater modeling to forecast hospital admissions in the US: Challenges and opportunities](https://community.epinowcast.org/t/community-seminar-2024-08-07-kaitlyn-johnson-wastewater-modeling-to-forecast-hospital-admissions-in-the-us-challenges-and-opportunities/237/17):
>
> (including variable names!) to work on the interface functions… which is really bad.

Agree this is a big issue though I am not sure how your suggestion helps lower the barriers for entry as new contributors would still not be able to add model features (as they would just be adding UI for those features?).

> [@Community Seminar 2024-08-07 - Kaitlyn Johnson - Wastewater modeling to forecast hospital admissions in the US: Challenges and opportunities](https://community.epinowcast.org/t/community-seminar-2024-08-07-kaitlyn-johnson-wastewater-modeling-to-forecast-hospital-admissions-in-the-us-challenges-and-opportunities/237/17):
>
> I guess that in `Julia` you could ideally take such a representation and directly construct the respective EpiAware model during runtime. And if someone wants to try build that with `stan`/`brms`, we won’t stop them

Yes exactly. I think this would be great if someone did but as below so much effort (see below)

I agree a `brms` like model generator could be the way to go but to be honest I am not sure a PPL agnostic front-end is less work/learning than just making people learn a new language (i.e. our Julia project 😉 ).

Do you have an example from other fields of the kind of domain specific tool you are thinking about existing?

In my head what you are proposing is a R specific PPL that then under the hood maps to other PPls and contains epi specific functionality. That sound very very hard and high effort to me? What happens if someone just does this mapping from say stan to whatever the new hotness is?

I think the version of this that remains focussed on stan (maybe only for now) is very doable though.

> (it remains the responsibility of the developer to implement the model correctly and to reject representations that your stan model cannot implement).

I think this is a separate issue and highlights the problem that I think most of the issue lies in developer time/skill as people (myself included) don’t implement the details correctly. If we had a really nice user interface across lots of still wrong packages that seems like a bad use of effort?

Given that most/all of the currently available tools that exist has fundamental flaws to their infra it seems like we should fix those (by ideally pooling resource) before we worry hugely about a extremely consistent UI. Especially as currently many/most users are relatively specialised/motivated and so can navigate maybe clunky UIs?

---

<div class="post-metadata">

**Author:** ![adrianlison](https://dub1.discourse-cdn.com/flex005/user_avatar/community.epinowcast.org/adrianlison/32/27_2.png) [@adrianlison](https://community.epinowcast.org/u/adrianlison)\
**Post date:** [14 August 2024 16:41 UTC](https://community.epinowcast.org/t/streamlining-of-epi-modeling-tools/241/13 "2024-08-14T16:41:26Z")

</div>

Thanks @samabbott, I agree that building a new PPL in R would be both very difficult and probably a waste of resources.

My original thought was not about having a complete PPL (although my remark about Julia seems to suggest this), but to simply have a structure in R to represent the inputs for the inference tool (in `stan`, this would be data and inits) which reflects the components of our models in a streamlined way. So I’m basically thinking of

tool-specific UI → intermediate data structure → flat list of `stan` arguments → `stan`

We partly have this “intermediate data structure” implicitly in our tools (e.g. modules in `epinowcast`), but I was thinking that if there was an explicit and standardized structure it could be

1. easier to handle (e.g. there could be options to print the inputs in a conceptually meaningful way, and one would better know what parts of the data must be modified when updating a UI function)
2. easier to map these inputs out to other inference tools than `stan`, i.e. requiring less refactoring of the tools-specific interface

For an example, think of a standardized structure in `R` to represent a renewal process, with certain agreed on attributes that map to variable names in your generative model, and nested attributes for non-parametric smoothing of the growth rate, for a seeding process, and for things suggested above (prior specification, formulas etc.). However this would not have to contain all the details of the modeled likelihood!

I guess my main point is that I think establishing something like the above would be roughly comparable in terms of effort and complexity to building a shared stan library, but potentially more useful in the long-term…
