Evaluation

Please note: A ParEvo exercise has two stages: generative and evaluative. This page is about the evaluation stage.The planning and implementation stage is described in a separate page here. Both are important. They are equivalent to the Exploration and Exploitation phases of organisational learning, as described by March.

1. During an exercise

1.1 Using Tags

When participants write their contributions they can be enabled to use one or more predefined tags to describe the content of thier contributions. The default tags are: Likely, Unlikely, Desirable, Undesirable. But the Facilitator can remove these and or add others, via the Edit Exercise page..

1.2 Using Comments

Participants can be enabled to post Comments on any of the contributions that have been made. The use of the Comment facility is optional. The Facilitator decides if and when to allow Comments and individual participants can choose if and when to make comments during any given iteration. They can make a maximum of one comment per storyline in a given iteration.

This comment facility can be used as the second phase of each iteration of the ParEvo process. But only after the window for contributions has been closed and all are now visible to all. Displayed comments are anonymised and can be seen, at the start of the contribution phase of the next iteration. Third parties, such as those with expert knowledge, can also be enabled as Commenters

1.3 Using the Evaluation widget

After the last iteration the Facilitator can trigger an Evaluation stage, where participants are asked to rate each of the surviving storylines, on two default criteria: (a) their probability of happening in real life, (b) and their desirability of happening. The Facilitator can also edit and change the default evaluation criteria. Other criteria, such as novelty, observability, sustainability or equity may be more useful in some circumstances.

After all responses have been received, participants responses to the widget’s evaluation questions can be made visible to all participants, through a display underneath the tree diagram on the left side. The ratings for each storyline can be viewed by clicking on the final section of each storyline branch.

2. After an exercise

2.1 Using an online survey

Online survey instruments, such as Survey Monkey, can be used to ask participants additional and more open-ended evaluation questions. The following types of questions can be asked:

  • Questions about specific storylines
    • Closed ended questions, like those above about desirability and probability
    • Open ended questions, such as the “most significant difference between the storylines”
  • Questions about the whole set of storylines
    • What was most surprising about the content, or what was absent from the content
    • How much the events were likely to affect the participant, and how much the participant could affect the events
    • Questions about participant’s judgements versus that of others
    • How optimistic their own contributions were, versus those of their own

There is a now a dedicated web page with a growing menu of different types of evaluation questions. You can also see an example of an external survey used in association with the Alternate futures for the USA 2020+ exercise

2.2 Using downloadable data

At the end of a ParEvo exercise, the following kinds of data can be downloaded in an Excel file format, and subject to further analysis:

  • Text of each storyline — organised by storyline
  • Text of each contribution, along with author and ID number
  • List of all comments
  • List of all participants
  • Evaluation matrix — showing which storylines were rates as most/ least desirable/probable
  • Tags x Contributions matrix
  • Tags x Storylines matrix
  • Participant x Participant matrix — showing who built on whose previous contributions
  • Participant x Storyline matrix — showing who built on which storylines
  • Participants evaluation judgements — showing which participant made which judgements
  • Export search history — showing which key words were searched for, their aggregate incidence and were found in which specific contributions

2.3 Debriefing meetings

It can be useful to hold post-exercise debriefing meetings (online or face to face) with all participants. There will necessarily be a loss of anonymity of participants at this stage, but anonymity of contributions can be maintained, if needed. For example, if the event was being used to gather more comparative judgments about different contributions (or events they describe).

3. Analysis options

3.1 Content analysis

All the methods described below involve choosing categories of storylines of interest, finding their members, and identifying significant differences between these. All with the ultimate intention of identifying implications for how participants should think and act in the future.

3.1.1 Missing storylines

Participants can be asked via the online survey, or in a workshop (online or face to face). What contents are surprisingly absent, given the focus of the exercise

Facilitators can use the Search field to see if any key words of specific interest are mentioned at all

They can also use LLM AI to identify other more complex-to-describe contents are surprisingly absent, given the focus of the exercise

Third parties can also be involved, via an Evaluation survey or the use of the Comment facility

3.1.2 Extinct versus surviving storylines

With each additional iteration the proportion of storylines that are extinct will grow. As soon as they approach a majority it makes sense to examine whether there are any significant differences between these and the surviving storylines, which might have consequences that need to be considered.

Methods for identifying these differences are essentially the same as with missing storylines: participant surveys and workshops, facilitator key word searches and LLM AI inquiries.

3.1.3 Differences between surviving storylines

These can be identified by two types of methods. Some useful differences will be predefined, because they are already part of a background theory a facilitator is bringing to the exercise. Some of these will be widely applicable across many exercises. Such as the likelihood and desirability of the storyline events. These are currently the default criteria built into the ParEvo evaluation widget. Others, such as equity or sustainability, might be more relevant to some exercises but not others.

The second type are not predefined, but are more inductive and emergent, born out of a comparison of the contents of the two types of storylines being compared. These differences can be sourced from participants (via surveys and workshop), facilitators, or LLM Ais. A suggested default question format to use is “Please compare these ….storylines and identify what you think is the most significant difference between them, in terms of their implications. We are especially interested in implications for (…fill in exercise context relevant text here…]. Describe the difference and the difference that might make.

3.1.4 Analytic frameworks

Having identified relevant differences, and the storylines that belong to each binary category created by the difference, there are four ways of organising this information that will help draw out implications for follow up actions.

3.1.4.1 A matrix view

Two difference criteria can be combined in the form of a two-by-two table, to highlight the four different types of storylines that might be found. The matrix below uses the two default criteria of likelihood and desirability. Within each cell there are some generically applicable responses that will be applicable to that combination, as shown below. Other combinations will probably require different generic responses.

Survey responses will have identified which storylines belong in each cell. The next task is to identify how the cell’s generic response could be articulated into more specific steps that could be implemented given the particular types of futures now located in each cell.

Either during this process, or separately, attention also need to be given to potentially useful generalist responses that might be appropriate across two or more cells in the matrix. That is, suitable to both types of likelihood, or both types of desirability, or all four cells

3.1.4.2 A tree view

After two initial groups of storylines have been identified by asking the “most significant difference” (MSD) question, that question can then be reiterated, focusing on each of those groups in turn. And then reiterated again, focusing on each of the two subgroups that are generated, and so on. This will generate an inverted tree structure like the one below (ignore the text contents).

This tree structure is a nested classification of the full set of storylines. It provides a multi level analysis of the alternative futures. At the top is high level description of the storylines in the two biggest groups. Further down are more granular descriptions of smaller sub-sets of the futures. At the bottom, in green, are the futures themselves.

At each junction a range of different evaluative and planning questions can be asked about those two kinds of futures. These fall into two types:

  • Differences in degree (more vs less). Such as
    • Likelihood of occurrence
    • Desirability of occurrence
    • Breadth of impact
    • Depth of impact
    • Avoidability
    • Achievability (technical and / or political)
    • Relevance to the participants lives or work
    • Ability of participants to influence
  • Differences in kind (categorical)
    • In what ways is this group more (e.g. one of the above comparisons) than the other?) than the other?
    • What is the most significant difference between these two groups in terms of their expected impact?
    • How we should react differently to these two types of these events?

A major adantage of this type of analysis is that it generates a multi-level analysis, with varying degrees of coverage and specificity (analagous to recall and precision). At the upper levels will be judgements with wide coverage, but with possibly a low fit to any specific future. At the lower levels will be judgements with narrow coverage but with better to fit to more specific futures. In others, words, an potential assortment of generalist and specialist strategies.

3.1.4.3 Scatter plot view

This is a way of comparing all branches, rather than pairs of branches. To start with, answers to each of binary (more vs less) question, at each branching point in the tree, are converted a binary score for each storyline category on each question. The two sub-groups are given a “1” (=more) or “0” (=less) as their answer to a comparison question. Looking down the entire length of any branch, that storyline could end up a with an aggregated binary score made up of a string of these judgements e.g. 11 , where there are two junctions (as in the toy example below). These binary scores can then be converted into numerical (base10) values i.e. 3 in this example. This is useful in two ways: (a) It provides a more holist view of the distribution of different kind of futures across the whole tree, not just one pair of categories at any given level. (b) These aggregated scores, for all storylines, on two questions of interest, can be displayed in a 2D scatter plot. to see how the four possible types of futures are distrubuted. And then to identify different types of possible responses for these.

3.1.4.4 Explotation versus Exploration

This is a distinction between two contrasting learning/survival stategies, widely publicised via James March’s 1991 paper . Claude AI has explained this distinction as follows:

March distinguished exploitation—refining and improving existing capabilities through efficiency and incremental refinement—from exploration—searching for new possibilities through experimentation and risk-taking. Exploitation yields more certain, near-term returns but can lead to competency traps, while exploration is essential for adaptation but involves uncertain outcomes and distant payoffs. Organizations face a fundamental tension in allocating resources between these strategies, as the reliable returns from exploitation tend to systematically drive out the uncertain investments needed for exploration.

In the context of a ParEvo exercise the first image (on the left) shown below represents exploration: it generates maximum variety by developing many divergent branches that evolve independently from each other, creating a wide range of distinct alternative futures without convergence.. The second image (on the right) represents exploitation: it shows repeated convergence around shared nodes (highlighted in yellow) with branches that develop more similar pathways, essentially refining and building upon common themes rather than maintaining maximum differentiation. These are extreme examples, in most ParEvo exercises the tree structure will be a more hybrid version.

The existence of as bias towards one of these two strategies can be identified by observing the number of original (iteration 1) storylines are still being develoloped in most current iteration. Which of these two bias that are preferred (or neither) will depend on the exercise context, especially its purpose.

3.2 Participation analysis

3.2.1 Why do so?

  • To identify ways in which the contents of storylines can be improved
  • To identify how objectives relating to peoples participation can be better achieved

3.2.2 What to analyse

3.2.2.1 Dropouts

In most exercises seen to date a small number of participants have dropped out during the exercise. When this happens these questions could be asked:

  • How did those that dropped out differ from the remaining participants?
  • Did they do so because of any dissatisfaction with the process?
  • Did their leaving have any influence on the subsequent storylines?

3.2.2.2 Contribution behaviour

When a participant makes a new contribution to existing storylines, they make two types of connections: (a) they connect to another participant — the one whose contribution they are immediately adding to, and in doing so (b) they connect to a specific storyline — a string of contributions that have been built on, one after the other.

In the Download section of the Edit Exercise page data on these relationships can be downloaded as two Excel files:

  • Relationships with other participants (known as an adjacency matrix)
  • Relationships with storylines (known as an affiliation matrix)

Both can be analysed using a range of social network analysis methods, available in software such as Ucinet/Netdraw

3.2.2.2.1 Participant relationships with each other

A scatter plot view

The simplest analysis is to calculate, for each participant, these two measures:

  • InDegree: The number of other participants who have built in their contributions
  • OutDegree: The number of other participants whose contributions this participant has built on

Then to create a scatter plot showing the distribution of values of participants on both measures, of the kind shown below.

The significance of this distribution will depend on the facilitator’s prior beliefs about participants behaviours and any expected ideal distribution. The design of the ParEvo process is not based on any assumption about the distributions of the four behaviour types, other than perhaps an implicit view that a diversity of behaviours might be of more value than the dominance of any one type

A whole network view

The same adjacency matrix data, when imported into software such as Ucinet/Netdraw, can be used to generate a network diagram, showing which participant was connected to which other participant, by giving or receiving contributions.

Two extremes are possible. One is a very densely connected network, where everyone is connected to everyone else. The other is a very sparse network, where each participant is connected to very few others. From a viewpoint of creativity as combinatorial process, a dense network structure suggests a much higher degree of diversity of ideas might arise because of the greater diversity of interactions. On the other hand, a ParEvo exercise on the UK Brexit experience suggests that a very sparse network is associated with more polarisation of views.

These generalizations should be treated with caution. As above, the significance of participants relationships network will probably best be interpreted in the light of each exercise’s unique context and the facilitators expectations.

3.2.2.2.2 Participants relationships to storylines

Three facets can be investigated:

  • Which surviving storylines are most collectively owned, in the sense that all participants have added at least on contribution to the storyline
  • Are any participants views being significantly marginalised? As evident in their contributions being noticeably more common in the extinct storylines?
  • Which participant has contributed the most to the surviving storylines, as a percentage of all their contributions?

3.2.2.3 Perceptions of the ParEvo Process

Process issues to investigate via a survey or workshop:

  • How important it was to participants that others built on their contributions
  • How important tit was to participants that they built on other participants contributions
  • On average, how many minutes they spent during a single iteration reading, then adding their own contribution
  • Was this longer or shorter than they initially expected
  • To what extend did they enjoy/dnot enjoy reading the storylines and then adding their contributions to the storylines
  • How useful were the Comments under the different contributions
  • Which Comment did they find the most useful, and why so
 Save as PDF