Friday, 20 December 2019

Audience Feedback from Presentations given by the DataViz Working Party


During 2019, the DataViz working party have given several talks to various IFoA audiences. These built on the presentation that the working party gave at the IFoA Life Conference in Liverpool in November 2018. We thought it helpful to summarise the audience feedback that we have received and the types of questions that have been asked.


We have presented at meetings and gatherings of actuarial societies in Scotland (Stirling, Glasgow, Edinburgh), England (Leeds, Leicester), and Switzerland (Zurich) - to approx. 200 people in total.


Presenting the aim and scope of our Working Party to fellow actuaries has been a great opportunity and rewarding experience for us. We have particularly enjoyed the multitude of interesting and relevant comments and questions from the audience, and would like to share and discuss these in this blog post.


The comments/questions have been structured into three categories:

  1. more general comments/questions related to the format or output of the Working Party or the interaction of the WP with other bodies;
  2. specific visualisation-related comments/questions, which often came in the form of suggestions for a future blog post on a particular topic;
  3. comments on already published blog posts.

1. General comments/questions about the Working Party

Actuarial Syllabus We've had a comment that data visualisation should be part of the actuarial syllabus, and the question whether we will try to push for this.
Collaboration A person has asked what engagement we have had with other members of the profession, or actuarial firms. We were also ask to liaise with other practitioners as actuaries don't have all the answers. We are in a regular exchange with the profession. We will soon also post a guest contribution from an American actuary who is not a member of the WP. Note that our blog is public, and we are happy to discuss questions and problems with anyone interested.
Promoting the blog We were asked to promote our working party via Facebook or LinkedIn. Output of the working party - We were asked what the final output of the working party will be. Our main output is the blog, and if everything goes as planned, the content will eventually be transferred to the profession's website.
Supporting the profession We were asked if we could help the profession with its public publications, e.g. communicating mortality improvements to the public.
User feedback Somebody asked how we would know that our examples are "better", i.e. whether we have asked for user feedback.
User support It was suggested that we could allow users to upload a dataset that they want to visualise but are having trouble with.

2. Visualisation-related comments/questions

General comment The questions below contain many good suggestions for future blog posts. We will try to cover at least some of these topics in the coming months.
Animations We have been asked by various people whether we could add a blog post on animated graphs. This is planned, as we also recognize this as something increasingly important
Visualisation software Multiple people have asked to share information on the use of R, Tableau or Power BI. We have already produced blog posts about visualisations created with R, and we will try to comment on the use of Tableau and Power BI.
Visualising stochastic results We have been asked multiple times to add a blog post on the presentation of stochastic results. This is in the pipeline, and we hope to have a blog post published relatively soon.
Visualisation of maps A person has commented that a blog post about geographic-mapping data would be helpful.
Reliability of data/analysis Someone has asked us whether we could comment on the reliability of the data/analysis/conclusions being communicated.
Good and bad practice It was suggested that we share our thoughts on good and bad examples of visualisation that we commonly see in actuarial work.
Research into thought processes Someone has asked us whether we are aware of relevant research into thought process / how the mind reads and analyses information that supports the kind of principles we are suggesting in our blog posts.
Well-known vs. novel visualisation types We have been asked to comment on the importance of considering what the reader is used to - a simple chart that a reader is used to seeing, compared to an innovate chart that the reader isn't used to seeing.
Data labels It was suggested that we comment on the use of data value labels in charts.
Colour schemes It was commented that red/green colour schemes don't print well in black and white, and that we should advise on other possible colour schemes (e.g. red-blue).
Blind users Someone asked how one could help blind users to understand a chart.
Questions one didn't have Good visualisations can sometimes answer questions one didn't have, and we have been asked to give an example for this.
Other visual clues Somebody has commented on the effectiveness of visual clues beyond colour schemes. This is a topic addressed in the 2014 SIAS paper of Julian and Paul (see Resources section of the Blog).

3. Comments on published blog posts

Client recommendations post It was commented that far too many recommendations are shown, and that the example would work better with c. 15
Mortality improvements post It was commented that the colour scheme was not ideal, and that we should consider CMI's best practice advice on heatmaps.
Correlation assumptions post Someone commented that the example is very complex at start and still complex at the end. It would be helpful to have an example where we distil something complex down to something much more simple.

Blog posts: past, present and future


In the spirit of looking back over the year and ahead to the next, this short post summarises our main posts over 2019 and looks ahead to 2020.

The past

During 2019 we've posted on:
  • Expanded hints and tips
  • Principles of Data Visualisation
  • Creating a Waterfall Chart in Excel
  • Data visualisation accounts on Twitter
  • How to Improve tables
  • Visualising Model Dependencies


The present

See this month's post on the feedback and questions that we've had from audiences at several talks we've given during 2019


The future

Upcoming posts in early 2020 are planned for:

  • Use of Business Intelligence Tools
  • A couple of posts on Visualising higher-order datasets
  • Visualising stochastic results
  • General insurance triangles


Sunday, 24 November 2019

Visualising model dependencies



1. Problem statement


When checking an actuarial model, common approaches are to do a detailed check of code and inputs (which should, in theory, pick up any errors in the model, but with the disadvantage that you can miss seeing the wood for the trees) and to test the reasonableness of outputs (with the advantage that this focusses on the implications of the model results for the real-world situation, but which might miss flaws in the details that would become critical under different parameterisations or uses).


We propose a way in which model parameters (and, in particular, the dependencies between them) can be visualised in order to provide a higher-level review of a model’s structure, to complement the above checks.



The example used here is a pension scheme valuation model, and in particular reviewing the programming of a pension scheme’s benefit structure.

2. Suggested approach



We suggest the use of a network map analysis of a pension scheme’s programmed benefit structure.


This was programmed as an interactive map so that the reviewer can vary the level of detail (for example, to zoom in on particular areas) given the large number of variables. The example output used in this blog post takes the form of screenshots from the interactive version.


Figure 1 below shows an example, focussing in on the calculation of accrued pension at normal pension age (which is 60 in this case) for one section of a large pension scheme.





This shows that, for example, one of the factors affecting accrued pension (the large purple dot, or Accrued NPA60 Pension) is the Final Pensionable Earnings (which is defined in the scheme rules) at the member’s future retirement. In turn, this will depend on the member’s full-time equivalent (FTE) salary at retirement, their date of leaving (DOL), and so on.

A pension scheme valuation model needs to reflect the wide range of benefits that a pension scheme might provide (for example, retirement pensions on early retirement, normal retirement or late retirement; any lump sum benefits paid on retirement or death; any ill-health retirement pension that might become payable; any deferred pensions payable where a member has changed jobs or otherwise left the scheme; any benefits paid to a spouse or other partner, etc). In each case, the scheme rules will specify the calculation of the member’s benefit, which will depend on a number of different elements. The network map allows all such parameters and dependencies to be visualised.

3. Rationale and commentary


Network maps show the connections (links) between different items (or nodes). Network maps can be either undirected (where the links show connections between the items with no direction specified) or directed (where one-way or two-way arrows indicate the directions of the connections). We have used a directed network map in this case.

The rationale for using a network map for this purpose includes:

  • Visualising the parameters in this way can check for consistency between the benefit programming of different schemes – do the network maps for similar schemes have a similar structure?
  • An initial review of the network map can highlight any unexpected features, for example a member’s retirement pension that depends on any assumed characteristics of their spouse or partner would usually be an error.
  • The network map easily highlights areas of complexity and simplicity, which might suggest further investigation is required.  For example, if a particular variable has an unusually high number of dependencies, this warrants a closer look.
  • Such a review of the model complements other checks, and might highlight features which might have been missed otherwise.
  • The network map allows for an easier view of the interdependencies of benefit scheme structures and could lead to a streamlining of the setup.

Here are a few examples of how the network map can be used:

Figure 2 shows the calculation of accrued pension for a scheme where some of the benefits have accrued with a Normal Pension Age (NPA) of 60. It sets out the items which it relies on – i.e. service and salary. The network graph clearly shows that the accrued NPA60 pension is correctly calculated using NPA60 service. However, in a previous iteration the graph helped to identify that the pension was incorrectly being calculated using some NPA65 service. 



Figures 3 and 4 below show an example with an unusually complex benefit structure. Checking the setup using the traditional format (which involves numerous lines of data interacting with each other) would take significant time and could still result in errors being overlooked. The interactive map allows for easier and more accurate identification of interactions between items. 

Figure 3 shows the large number of items which are linked to a particular benefit item, and Figure 4 shows the same item with a number of the more distance nodes removed. This can be enabled via a distance selector in the tool. This feature helps the user to focus on particular parts of the structure. 






Figure 5 below highlights an additional feature of the tool which allows the network map to be shown in various formats. Figure 5 shows the same benefit item as shown in figure 3 but the colour coding highlights the different states at which the calculations are carried out, in particular distinguishing between the calculations that are carried out at the point at which the member joins the scheme, to withdrawal, to retirement and finally death of the pensioner. This feature allows the user to concentrate on specific areas which might be causing an issue in the model.







4. Applicability and alternatives



In theory such an approach can be used with any model. Its applicability will depend on the complexity of the model and its interaction with the suite of other model checks.


5. Implementation


The example illustrated here was developed using R studio, also using Shiny (an R package which makes it easy to build interactive web apps).


The R code reads a Microsoft Excel file which contains the detailed benefit structure including the interdependencies between data and benefit items. The code processes the data in the Excel file to create tables of the nodes and the node connections by assigning a unique identifier to each node.



The tables are then processed to produce the necessary nodes and connections to display in the interactive network map.


6. Context


Pension schemes provide a wide range of benefits to members and their dependants. The benefits payable will depend on factors including:

 
  1. What happens to the member (and their dependant in the future), for example whether they are unable to work due to ill-health, whether they choose to take their retirement pension after their normal retirement date and how long they live for; 
  2. The member’s work and salary history (including assumed future experience where the member is still in service); 
  3. The precise benefit specification (which is set out in the scheme rules – sometimes also affected by legislation); and
  4. External variables (eg the rate of consumer price inflation)

 


Many of these features interact.  In order to value the liabilities of a pension scheme (which is usually carried out in order to assess the sufficiency of the assets held in the scheme), a pension valuation model combines information on scheme members with detail of scheme benefits and a number of assumptions about unknown future experience (such as longevity and inflation assumptions).



The approach described above complements other checks on the model’s programming and parameterisation in order to reduce the risk of model and parameter error.