19 - 527_Rating_Scales
- 01 - 1. General principles
- 02 - 2. Diagnostic schedules
- 03 - 3. Depression rating scales
- 04 - 4. Alcohol rating scales
- 05 - 5. Scales used in child psychiatry
- 06 - 6. Scales used in old age psychiatry
- 07 - 7. Other clinical rating scales
- 08 - 8. Outcome scales in psychiatry
- 09 - Some examples of outcome scales
- 10 - 9. Psychometry of rating scales
- 11 - 10. Risk assessment
- 12 - Approaches to risk assessment
- 13 - Clinical approach
- 14 - Actuarial approach
- 15 - Structured professional judgment
- 16 - Stages in risk assessment
- 17 - Commonly used tools
- 18 - Structured Risk Tools
- 19 - Actuarial instruments
01 - 1. General principles
1. General principles
© SPMM Course
- General principles Rating scales give clinicians an objective benchmark to support critical treatment decisions. Regular use of rating scales can provide information that aids in diagnosis, prognosis and therapeutic monitoring. Further, when self-report scales are used, an item-by-item analysis can help to identify the exact symptoms that a patient considers most troublesome and challenging so they can be targeted during the consultation process. Rating scales can be used for (1) screening for the presence of a psychiatric condition (2) diagnosis of a psychiatric illness (such scales are often termed diagnostic schedules) (3) estimating severity of various conditions and their response to treatment (4) assess functional capacity and well-being. Rating scales can be either self-rated or observer rated. Some observer-rated scales require clinical experience (clinician-rated) while trained non-clinical personnel can use others. Self-rated Observer rated Beck’s Depression Inventory Zung Depression Inventory Symptoms checklist SCL GHQ Lunser’s (extrapyramidal) Edinburgh postnatal depression scale Dementia scales PANSS BPRS HAMD MADRS YBOCS SCID Scales are based on psychometric properties; aim to measure dimensions of psychopathology (symptoms) often at the present state (the duration of which is defined variously). Schedules are based on clinical expectations; deal with categories of disorders (syndromes) based on known classification systems. A schedule may be devised to focus on either past or present status or both. Two types of procedures are used when selecting items (symptoms) for scales. The first method is based on the previous clinical literature to determine the appropriate items. The second method is based on calibration. In this method, a large number of questions are tested to find the most discriminating items between a ‘known ill’ and a ‘presumed well’ group. The items that are most significant in a statistical sense are chosen to represent the scale that is devised. Considerations for selecting a screening measure for use in a study include Characteristics of the population to be screened, Psychometric properties of the instrument, Time required to complete the measure, Ease of use, and Cost of obtaining the measure
© SPMM Course General Health Questionnaire is an all-purpose screening tool that is often used as a first-level assessment instrument in epidemiological studies before detailed diagnostic schedules are employed. Goldberg introduced the General Health Questionnaire (GHQ) GHQ was developed as a screening tool to detect those likely to have or be at risk of developing psychiatric disorders It is available in a variety of versions using 12, 28, 30 or 60 items, the 28-item version is used most widely. Each item is scored as a 4 point Likert (0-3) allowing a total possible score on the GHQ 28 of 0 to 84. Using alternative binary scoring method (with least symptomatic items scoring 0 and the most symptomatic items scoring 1), the 28- and 30-item versions classify any score exceeding the threshold value of 4 as achieving ‘psychiatric caseness’. The caseness threshold is 3 for the 12-item version. Psychiatric caseness is a probabilistic term—whereby, if such respondents presented in general practice, they would be likely to receive further attention. It should be noted that the GHQ is not usually used for predictive purposes. Reliability coefficients have ranged from 0.78 to 0.95 in various studies.
02 - 2. Diagnostic schedules
2. Diagnostic schedules
© SPMM Course 2. Diagnostic schedules
Scale Mode of administration Features Clinical Interview Schedule (CIS)
Clinician administered – fully structured. A revised version for lay interviewers also available (CIS-R). Used in National Psychiatric Morbidity Surveys of Great Britain. Developed by Goldberg et al. Aims to identify common disorders found in primary care & community settings (focus on neurotic conditions). Composite International Diagnostic Interview (CIDI) Clinician administered CIDI was an improvised schedule that incorporated principles of both PSE and DIS. It was produced by WHO to be used with both ICD and DSM diagnoses. Diagnostic Interview Schedule (DIS) Non-clinicianadministered; fully structured interview. Used in ECA study. Lifetime DSM diagnoses made initially; later a time period can be specified for ‘current diagnoses’. Hopkins Symptom Check List (HSCL) Trained primary care workers (HSCL-25) or self-reported (original) HSCL has a 58 items self-report version that measures ‘neurotic’ symptom distress in outpatients (somatisation, OCD symptoms, interpersonal sensitivity, anxiety and depression) and a 25 item objective version that measures symptoms of anxiety (10 items) and depression (15 items). SCL-90R and Brief Symptom Inventory (BSI) are derivatives of HSCL. Patient Health Questionnaire (PHQ)
Self-report scale It is the self-report version of Primary Care Evaluation of Mental Disorders (PRIME-MD) developed by Spitzer et al. on the basis of DSM-III. Aims to diagnose common neurotic conditions in primary care settings. PHQ-9 is a derivative that focuses on the 9 depression criteria in DSM-IV. GAD-7 on anxiety and PHQ-15 on depression with somatic features. Present State Examination (PSE)
Clinician-administered semi-structured clinical interview Provides clinical diagnoses in the lines of ICD system. CATEGO is the computerize version of PSE schedule. Schedule for Affective Disorders and Schizophrenia (SADS) Clinician-administered semi-structured Covers all major mental illnesses (depression, bipolar disorder, schizophrenia and anxiety disorders). A regular, a lifetime and a change version are available. A children version called Kiddie-SADS is also available. Schedule for Assessment in Neuropsychiatry (SCAN) Semi-structured interview for use by trained clinicians Developed by Wing et al. on the basis of PSE and has replaced PSE at present. Focused on adult psychopathology. Extensive (28 sections and 1872 items in total, but many can be skipped) Structured ClinicalInterview for DSMIV (SCID) Clinicianadministered semistructured To be used for patients in whom a psychiatric diagnosis is suspected. Non-patient version available for epidemiological studies. SCID-II is available for axis 2 disorders.
© SPMM Course Some general observations on diagnostic schedules currently available in psychiatry
- Most schedules are either DSM or ICD based; only a few cater both simultaneously.
- A number of primary-care oriented schedules focus largely on nonpsychotic disorders
- Many schedules have computerized forms; nevertheless a number of them are timeconsuming to complete routinely or during initial clinical contact.
- Almost all of them are clinician-administered though self-report forms have evolved in recent times.
03 - 3. Depression rating scales
3. Depression rating scales:
© SPMM Course 3. Depression rating scales: 2 Questions scale (also called PHQ-2): An affirmative response to the following two questions may be as effective as using longer screening measures or may indicate the need for the use of more in-depth diagnostic tools: (1) "Over the past two weeks, have you ever felt down, depressed, or hopeless?" and (2) "Have you felt little interest or pleasure in doing things?" HAMD/HDRS – Hamilton depression rating scale: Observer rated. 17 – 21 items- 2 versions Refers to last 1-2 weeks More items for biological features Remains a reference standard
MADRS – Montgomery-Asberg Depression rating scale: 10 items version Most sensitive to change Requires clinical interview like HDRS
BDI – Beck depression inventory: Self-rating 21 items Max score 63 Last 2 weeks profile 0-13 minimal; 14-19 mild; 20-28 moderate; >28 severe Lacks discriminatory power among very severely ill More psychological than somatic factors included Can be repeated at short intervals
Zung Depression Inventory: 20 items Self-rated Avoids imbalance towards psychological factors seen in Beck’s Poor correlation with observer rating Insensitive to change
Visual Analogue Scale (VAS) Easiest way to quantify depression severity 10cm line where patient indicates the state of mood
© SPMM Course Depression screening in special cases: Children with depression Depression screening for children and adolescents are generally appropriate in children who are at least seven years of age. Reynolds Child Depression Scale and the Children's Depression Inventory (full version with 27 items; screening version with 10 items) were developed specifically for children and are written at lower reading levels. Measure Age appropriateness (years) Children's Depression Inventory (CDI) 7 to 17 Center for Epidemiological Studies-Depression Scale for Children (CES-DC) 12 to 18 Center for Epidemiological Studies-Depression Scale (CES-D) 14 and older Reynolds Child Depression Scale 8 to 12 Reynolds Adolescent Depression Scale 13 to 18 Beck Depression Inventory (BDI) 14 and older Mood and Feelings Questionnaire (MFQ) 8 to 18 Adapted from an online version at www.aafp.org/afp/20020915/contents.html. Advantages of the BDI and CES-D include ease of scoring, low-cost, and comparable psychometric properties. MFQ is endorsed by the NICE and has a self and a parent-rated versions. Perinatal depression: The BDI, CES-D, and Edinburgh Postnatal Depression Scale (EPDS) have been used to screen for depression in women during the antepartum and postpartum periods. The BDI and CES-D tend to produce higher scores and more false-positive results in symptomatic pregnant women. Edinburgh Postnatal Depression Scale was specifically developed for assessing postpartum depression and relies much less on somatic questions. Questions on the Edinburgh scale (10 items, can be self or clinician-rated) are framed within the "past seven days", and the response format is frequency-based. Routine use of EPDS during the postpartum period has been shown to increase the detection of postpartum depression compared with usual care. Geriatric depression: The GDS – Geriatric depression scale was specifically developed for use in geriatric patients, and it contains fewer somatic items. Questions pertain to symptoms within the past week, and responses require only a "yes" or ‘no’. In patients who have cognitive deficits, interviewer-administered instruments such as the Cornell Scale for Depression in Dementia or the Hamilton Rating Scale for Depression are preferred.
© SPMM Course The Cornell measure should be administered to the patient's primary caregiver. The Brief Assessment Schedule Depression Cards (BASDEC) system is designed for general hospital use and eliminates the likelihood of questions being overheard on geriatric wards. Patients choose answers from a deck of 19 cards presented one at a time. Depression in schizophrenia: Most depression scales are tuned to assess depression in nonpsychotic patients. These scales contain items, which do not distinguish depressed from nondepressed psychotic patients (e.g. delusions of nihilism that can be present in psychosis in the absence of depression). Calgary Depression Scale for Schizophrenia (CDSS) focuses on symptoms of depression in the presence of schizophrenia.
04 - 4. Alcohol rating scales
4. Alcohol rating scales:
© SPMM Course 4. Alcohol rating scales: CAGE Questions: 1. Have you ever felt like cutting down on your drinking? 2. Have people annoyed or criticized you for drinking? 3. Have you ever felt bad or guilty about your drinking? 4. Have you ever had a drink first thing in the morning to steady your nerves or to get rid of a hangover (eye-opener)? A positive answer should raise suspicion of an alcohol problem, and a score of 2 is highly suggestive. The instrument takes less than a minute to administer. AUDIT: Alcohol Use Disorders Identification Test; Saunders et al., 1993: This is a 10-item questionnaire, covering quantity, frequency, inability to control drinking, withdrawal relief, loss of memory, injury and concern by others. A score of 8 or more indicates that the person is drinking to a degree that is harmful or hazardous, whereas a score of 13 or more in women and 15 or more in men is indicative of dependent drinking. It is widely used and recommended by WHO for primary care use. MAST: Michigan Alcohol Screening Test (Selzer, 1971): This is a simple, self-report, 25-item test, which has yes/no answers. A score of 3–5 is an early indicator of a problem drinker, whereas someone who scores 6 or more is highly likely to be a problem drinker. There are variants on this test, e.g. the Brief MAST, which can discriminate problem drinkers from nonproblem drinkers on the basis of 10 items only. There are also a G-MAST and a Brief G-MAST for older people with slightly different phraseology. CIWA: Clinical Institute Withdrawal Assessment for Alcohol: (Sullivan et al., 1989) Health professionals use this scale to rate the severity of alcohol withdrawal. It consists of 10 items, 9 of which can be scored in a range of 0–7 and one on a range of 0–4 (total of 67). It can be used regularly throughout the day and night to assess the extent of withdrawal and the impact of treatment. It covers nausea; tremor; paroxysmal sweating; anxiety; agitation; auditory, visual and tactile disturbances; headache and orientation.
05 - 5. Scales used in child psychiatry
5. Scales used in child psychiatry
© SPMM Course 5. Scales used in child psychiatry
The Child & Adolescent Functional Assessment Scale: A rating scale to assess the degree of impairment in functioning due to emotional, behavioural, or psychiatric problems. It is completed by a clinical staff and takes about 10 minutes. It is useful for assessing outcome over time and for directing case management activities. It measures aggression and conduct problems especially in age between 7 to 17 years.
The Child Behavior Checklist (CBCL): It records the behavioural problems and competencies of children aged 4 through 16, as reported by their parents or others (e.g., teachers) who know the child well. The checklist is composed of 113 items in a Likert scale. The instrument provides three scores: a total score and scores on internalizing behaviours (fearful, shy, anxious, and inhibited) and externalizing behaviours (Aggressive, antisocial, and under controlled). This instrument can either be self-administered or administered through an interview. The CBCL can also be used to measure a child's change in behavior over time or following a treatment. Teacher Report Forms, Youth Self-Reports and Direct Observation Forms are also available for the Child Behavior Checklist. Two versions of this instrument exist: one for children ages 1 1/2 - 5 and another for ages 6
Diagnostic Interview Schedule for Children (DISC): Originally developed in the early 1990s Fully structured diagnostic interview for making DSM-based diagnoses in children.
Conners Rating Scales A family of instruments that measure a range of childhood psychopathology Most commonly used in the assessment of ADHD. Teacher, parent, and self-report (for adolescents) versions are available Both short and long (up to 80 items) forms are available with
06 - 6. Scales used in old age psychiatry
6. Scales used in old age psychiatry
© SPMM Course 6. Scales used in old age psychiatry
The Geriatric Mental State Schedule (GMSS) is a widely used instrument for measuring a variety of psychopathology in community surveys of the elderly. AGECAT is a computerised algorithm based on GMSS. The Mini-Mental State Examination (MMSE) is a popular cognitive screening instrument for the elderly. It takes 10 minutes to administer by a trained interviewer. A cut-off score of 23 for the presence of cognitive impairment has been suggested. Educational status affects MMSE scores. MMSE does not pick frontal lobe deficits, a major drawback of using it as a screening instrument for dementias. It is also claimed to have only moderate to minimal sensitivity to change in mild cognitive impairment states. Abbreviated Mental Test Score is a brief 10-point questionnaire to assess memory and orientation in 3 minutes (Hodkinson, 1972). Its origins can be traced back to the Blessed Dementia Scale. A cut-off score of 7/8 out of 10 is suggested to suspect cognitive impairment in the elderly. The Alzheimer's Disease Assessment Scale (ADAS) is a standardised assessment of cognitive function, and non-cognitive features that take 45 minutes to be administered by a trained professional. The cognitive section is termed ADAS-Cog. It is the gold standard for measuring the change in cognitive function in anti-dementia drug trials. A fall of about 10% per year is expected (deemed average) in Alzheimer's disease. The BEHAVE—AD is a clinician-administered scale to document behavioural symptoms in patients with Alzheimer's disease. It covers paranoid and delusional ideation; hallucinations; activity disturbances; aggression; diurnal variation; mood; and anxieties and phobias. The Neuropsychiatric Inventory (NPI) can be used to record severity of associated behavioural symptoms of dementia over ten domains: (delusions; hallucinations; dysphoria; anxiety; agitation/aggression; euphoria; disinhibition; irritability/lability; apathy; and aberrant motor behaviour). It is scored from 1 to 144. The severity and frequency of behavioural symptoms are independently assessed. MOUSEPAD stands for Manchester and Oxford Universities Scale for the Psychopathological Assessment of Dementia. It is administered to carers by an experienced clinician for the measurement of behavioural and psychiatric symptoms of dementia (BPSD). Clifton assessment procedure for the elderly - CAPE (Pattie & Gilleard, 1979) is intended to assess the level of disability and estimate need for care in the elderly. It consists of a short cognitive scale and a behavioural rating scale. The latter has four sub-scales: physical disability, apathy, communication difficulties and social disturbance. It is quick and easy to administer.
© SPMM Course Bristol Activities of Daily Living Scale assesses 20 daily living abilities in patients with dementia. It is a caregiver-rated scale designed for community use by trained health professionals.
07 - 7. Other clinical rating scales
7. Other clinical rating scales
© SPMM Course 7. Other clinical rating scales SCALE Mode of administration Features Positive and Negative Symptom Scale (PANSS) Clinician-administered rating scale For assessment of severity and monitoring of change of symptoms in patients with a diagnosis of schizophrenia. 30 items are covering positive symptoms, negative symptoms, and general psychopathology. Yale-Brown ObsessiveCompulsive Scale (YBOCS) Clinician-administered semi-structured interview Allowing rating of severity in patients with a pre-existing diagnosis of OCD. SCOFF
SCOFF is a mnemonic for eating disorder screening (similar to CAGE for alcohol). It has a high sensitivity (2 or more questions positive).
Do you
- Make yourself SICK when you feel uncomfortably full?
- Worry you have lost CONTROL over how much you eat?
- Recently lost more than 14 pounds within three months?
- ONE stone's worth of weight
- Believe you are FAT when others say you are too thin?
- Would you say that FOOD dominates your life? Minnesota Multiphasic Personality Inventory (MMPI)
Results generate information useful for a broad range of clinical applications.
A self-report questionnaire consisting of 567 questions covering 8 areas of psychopathology, 2 additional areas of personality type, and 3 scales assessing truthfulness. Results are compared with normative data from non-clinical populations. NOT A PROJECTIVE TEST. International Personality Disorder Examination (IPDE)
Semi-structured clinical interview for use by clinicians producing ICD-10personality disorder diagnosis. 67 standardized probe questions. 57-item true/false questionnaire also included for screening purposes.
Clinical Global Improvement (CGI) Clinician rated based on clinical judgment A two-item instrument - CGI-S (severity) – the current condition on a scale of 1–7 & CGI-I (improvement) – the extent of improvement since the start of treatment on a scale of 1–7. Can be used for any psychiatric disorder encountered in a clinic or ward.
Brief Psychiatric Rating Scale (BPRS) Rated by the physician on the basis of a semi-structured interview (18 items, 7 points for each, maximum of 108) Developed by Overall, 1960. One of the most widely used clinical rating scales. It was originally intended for use in controlled clinical trials of new psychotropic drugs. However, it has also been widely employed in studies of the clinical (that is, symptom) correlates of cognitive and neurobiological phenomena. The ratings include observations as well as patient reports. Factor analysis yields five factors (hostility-suspiciousness, withdrawalretardation, thinking disturbance, depression-anxiety, and activation).
© SPMM Course Scale for Assessment of Positive Symptoms (SAPS) and the Scale for Assessment of Negative Symptoms (SANS) Clinician rated based on clinical interview Not intended as diagnostic devices. They have been used primarily in studies of the neurobiological correlates of symptom groupings. SAPS - 34 items divided into hallucinations, delusions, bizarre behavior, and formal thought disorder. SANS comprises 25 items divided into affective flattening or blunting, alogia, apathy, asociality, and inattention. Personality Diagnostic Questionnaire-4+ (PDQ4+) Self-report instrument assessing 12 personality disorders described in the DSMIV using 99 true-false items. A brief structured interview (Clinical Significance Scale) is used as a follow-up after the self-report to estimate whether (a) the trait is enduring (criterion D for DSM-IV); (b) it is present in the absence of other disorders (criteria E and F); and (c) it leads to distress or impairment (criterion C).
08 - 8. Outcome scales in psychiatry
8. Outcome scales in psychiatry
09 - Some examples of outcome scales
Some examples of outcome scales
© SPMM Course 8. Outcome scales in psychiatry The purpose of clinical intervention in psychiatry is to achieve a desirable outcome for the patient, his/her family and the society. Treatment outcomes are not single constructs but are multidimensional. Outcome measures can be broadly classified as follows: Psychopathological rating scale: measuring individual symptom severity in a disorder e.g. PANSS, BPRS. These scales can also provide modified measures such as relapse, remission, etc. that are sued as outcome variables in various studies. Global outcome measure: an overall appraisal of disease severity and its impact on overall functioning e.g. Global Assessment of functioning (GAF) or Global Assessment Scale (GAS). Generic patient-based outcome measure: measuring several domains of health-related quality of life applicable to populations irrespective of illness. e.g. Short Form 36 (SF36) Disease-specific patient based outcome measure: measuring several domains of healthrelated quality of life applicable to specific patient groups. Domain specific patient based outcome measure: measuring a specific domain associated with health-related quality of life e.g. focusing on social function or interpersonal relationships or cognitive capacity or service-level satisfaction. Any useful outcome measure must satisfy the following criteria to be incorporated into trials and clinical practice:
Some examples of outcome scales Global Assessment of Functioning Scale (GAF) is available as an appendix to DSM-IV. It measures overall psychosocial functioning in patients on a 10-item (100 points) scale, rated on the •Is the instrument appropriate to the question addressed by the trial or the benefit desired by the clincial service? Appropriate Appropriate •Is the measure reproducible, internally cosnsitent, precise and accurately reflecting the construct of interest? Reliable & valid Reliable & valid •Is it sensitive to change over time? Responsive Responsive •Are the scores intuitively meaningful and comparable to real life states? Readily Interpretable Readily Interpretable •Is the tool readily applicable in clincial settings and not too onerous on clinicans and patients? Acceptable & feasible Acceptable & feasible
© SPMM Course basis of self-report and information from the clinical interview. GAF was used in axis V of the multiaxial diagnostic system in DSM-IV. It combines symptomatic severity and functional impairment and is based on the clinician’s appraisal of the functional limitation. Social and Occupational Functioning Assessment Scales (SOFAS) was proposed as a new axis in Appendix B of DSM-IV. It is related to GAF but focuses only on functioning and not on symptoms. SOFAS does not try to discriminate between functional changes related to psychiatric and nonpsychiatric causes. It is rated by clinicians on a 100-point scale based on all available information, with descriptors for each 10-point interval. Guided tools for administering SOFAS (e.g. Personal and Social Performance scale –PSP) are available for use in clinical trials. Short Form health survey-36 or SF-36 was designed for wide use in a variety of settings: clinical practice, research, policy evaluations, and community surveys. SF-36 can be self-administered by persons 14 years of age and older, or by a trained interviewer. It assesses eight health concepts: 1) limitations in physical activities because of ill health; 2) limitations in social activities because of physical or emotional problems; 3) limitations in role performance due to physical health problems; 4) bodily pain; 5) general mental distress and well-being; 6) limitations in role performance because of emotional problems; 7) vitality (energy and fatigue); and 8) general health perceptions.
The health of the Nation Outcome Scales (HoNoS): HONOS is the most widely used psychiatric outcome scale within the NHS. It was developed by the RCPsych and commissioned by the Department of Health in 1993. It has 12 items measuring behaviour, impairment, symptoms and social functioning, measured on the basis of routine clinical assessments in various clinical settings and has influenced various policymaking processes in the English NHS over the last 2 decades.
10 - 9. Psychometry of rating scales
9. Psychometry of rating scales:
© SPMM Course 9. Psychometry of rating scales: When developing measurement scales, we are concerned about two important properties. Can we use this scale to measure the actual phenomenon we want to measure? Can this scale provide consistent results when it is used? A highly valid scale will measure what it is supposed to measure – the truth. A highly reliable scale will provide consistent results. Reliability refers to the replicable nature of research studies / tools. Note that high reliability does not guarantee scientific validity but guarantees consistency. Reliability can be assessed by test-retest correlation by administering an instrument twice to the same population. The time difference between test and retest must be long enough to avoid practice effect, but short enough so the underlying state (e.g. depression) does not change very much: 2 to 14 days range is often used in psychiatry. Cronbach’s alpha measures the internal consistency of a test by correlating each item with the total score and averaging the correlation coefficients. It can take values between negative infinity and 1 as a maximum; but only positive values make sense. Arbitrary cut-off of 0.70 is used commonly to call the evaluated test to be internally consistent. The split-half reliability refers to splitting a scale into two parts and examining the correlation. Interrater reliability is measured using two or more raters rating the same population using the same scale. The intraclass correlation coefficient is used for continuous variables; it is nothing but the proportion of total variance of the measurement that reflects true between subject variability. It ranges between 0 (unreliable) and 1 (perfect reliability). ICC can be measured by either relative agreement or absolute agreement; the relative ICC is always higher than the absolute ICC. ICC of 0.6 is considered fair while 0.8 is very good and 0.9 as excellent, arbitrarily. ANOVA intraclass coefficient is used for quantitative data with more than 2 raters/groups. For nominal data that has more than two categories, a kappa or weighted kappa can be used. (More details are given below) Validity of an instrument is the extent to which an instrument measures what it proposes to measure. Face validity refers to a subjective measure of deciding whether the test measures the construct of interest on its face value.e.g., Hamilton depression scale clearly has a face value in measuring depression; but not for measuring obsessions.
© SPMM Course Construct validity measures whether a test really measures the (theoretical) construct of interest or something else. One way of classifying the construct validity is considering unified construct validity. Here construct validity is taken to consist of both content validity and criterion validity (referred as unified construct validity). Content validity refers to whether the contents i.e. each individual subscales, items or elements of the test are in line with the general objectives or specifications the test was originally designed to measure. It looks for a good coverage of all domains thought to be related to the measured condition. This often cannot be statistically tested, but experts are called for comments on this aspect of validity. Criterion validity refers to the performance against an external criterion such as another instrument (concurrent) or future diagnostic possibility (predictive). Concurrent validity refers to the ability of a test to distinguish between subjects who differ concurrently in other measures (using other instruments). e.g., those who score high on a scale of insomnia may score high on a scale of fatigue ratings too. Predictive validity refers to the ability of a test to predict future group differences according to current group differences in score. e.g., high aggression score in childhood and high criminal incidents in adult life. (On a similar note, Incremental validity refers to the ability of a measure to predict or explain variance over and above other measures)
Another way of considering the construct validity is by classifying it to convergent, discriminant and experimental/interventional validity: Convergent validity refers to agreement between instruments that measure same construct e.g. between BDI and HAMD for depression. This agreement can be tested in contrasted groups i.e. depressed and non-depressed, both groups showing a high correlation between the two scales. Discriminant validity refers to the degree of disagreement between two scales measuring different constructs. e.g., to say that HAMD measures some construct (depression) different from that measured by Hamilton Anxiety scale (anxiety) poor correlation must be demonstrated between HAMD and HAS Experimental validity: This refers to the sensitivity to change. An instrument must show the difference in results when an intervention is carried out to modify the measured domain.
© SPMM Course Note: Factorial validity is a form of construct validity established via factor analysis of items in a scale. Precision and accuracy Precision is the degree to which a calculated central value (e.g. mean) varies with repeated sampling. The narrow the variation, the precise the value is. Random errors lead to imprecision. Factors reducing precision includes 1. Having wider the limits of the interval 2. Expecting higher confidence interval (e.g. 99.7% versus 95%). Accuracy refers to the correctness of the mean value – i.e. how close is it to the true population value. Precision is comparable to reliability while accuracy is comparable to validity. Bias in a study compromises validity / accuracy. VALIDITY QUESTION IT ANSWERS Face Does this scale appear to be fit for the purpose of measuring the variable of interest? Content Does this scale appear to include all the important domains of the measured attribute? Criterion Is the scale consistent with what we already know (concurrent) and what we expect (predictive)? Convergent Does this new scale associate with a different scale that measures a similar construct? Discriminant Does the new scale disagree with scales that measure unrelated constructs?
11 - 10. Risk assessment
10. Risk assessment
© SPMM Course 10. Risk assessment Risk is the likelihood that harm will occur. Risk assessment is a process of scientific (statistical/clinical) calculation of likelihood of an adverse event; this includes specifying
- What will happen?
- When will this happen?
- By whom will this happen?
- How will this happen? Risk management is integral to assessment; Assessment is not carried out to label people and categorise into groups – it is a dynamic process and specific to the event of importance. Risk cannot be eliminated but only reduced. Problems with ‘predicting’ risk: Low base rate: The events of interest are usually very rare. Hence the predictive value is generally low. Serious violence is rare amongst the severely mentally ill disordered, killing or maiming of others is measured in probabilities of less than 1% (Wallace 1998). This low base rate seriously compromises the predictive utility of risk assessment because of the false positive rate of even an exceptionally good risk assessment instrument. Multifactorial: Risk is dependent on several factors, which tend to change over time. Unknown interactions: Comprehensive risk evaluations are time-consuming; often the degree and nature of interaction among various factors is unknown. Risk factors for any untoward incident (suicide, crime or violence) can be categorized as static, stable and dynamic factors (Bouch & Marshall, 2003). Static risk factors: These are fixed and historical: e.g. family history of suicide. They cannot be modified. Stable risk factors: These are long term, enduring issues but are modifiable to some extent and not fixed: e.g. diagnosis of personality disorder. Dynamic risk factors fluctuate markedly in both duration and intensity: the e.g. presence of acute anxiety symptoms or akathisia. A dynamic risk factor can act synergistically and multiply the effect of underlying static and stable risk factors if not addressed promptly. Certain dynamic factors may occur only in future (e.g. upon discharge, a patient may feel helpless). A comprehensive risk assessment should consider static, stable, dynamic and future risk factors and include them in devising risk management strategies.
12 - Approaches to risk assessment
Approaches to risk assessment
13 - Clinical approach
Clinical approach
14 - Actuarial approach
Actuarial approach
15 - Structured professional judgment
Structured professional judgment
© SPMM Course Approaches to risk assessment Clinical approach In this approach, a clinicians’ subjective, intuitive judgment informed by experience and knowledge is used to estimate risk and guide decisions about treatment. But professional opinions are often highly variable for a given case and have poor predictive value, not supported by many policymakers. Most risk assessment currently performed by clinical methods. They are based on clinical experience, individuals knowledge and person-specific assessment to reach a conclusion. Only 1/3rd of pure clinical judgments are estimated to be correct in retrospective studies. Actuarial approach This approach is very popular in forensic services. This uses formal, algorithmic and objective procedures for quantifying risk as a numerical probability of a future outcome: e.g. patient A has a 60% chance of killing herself in the next 2 years, etc. It is found to be superior to other methods of predicting the risk of violence and sexual offending, but does not inform the clinician much about the risk factors that require targeting to mitigate risk. There are hardly any actuarial tools available for self-harm and suicide risk, but there are many available for violence risk assessment. Problems of actuarial tools include:
- Historical aspects are given more importance – so a pessimistic view of risk with insensitivity to change results.
- High false positive rates using these tools.
- As mentioned above generalisation of actuarial tools developed in one setting to the other is difficult.
- Further, actuarial approaches are often too focused on static and stable risk factors rather than dynamic and modifiable factors. Thus, though actuarial instruments more accurately identify at-risk groups, structured risk instruments may be more clinically useful for enhancing clinical decision-making. Structured professional judgment Structured professional judgment is an approach that aims to combine the evidence base for risk factors with an individual clinical assessment to complement psychiatric opinion. A structured, scales-based assessment is used when formulating risk management plan. Several risk assessment instruments are available to support structured professional judgment (e.g. HCR–20 to assess risk for violence in forensic settings) It is important to note that neither actuarial nor structured judgments should replace conventional clinical assessment; they must be employed as an additional aide for systematically identifying and addressing relevant risk factors.
16 - Stages in risk assessment
Stages in risk assessment
17 - Commonly used tools
Commonly used tools
18 - Structured Risk Tools
Structured Risk Tools
© SPMM Course Stages in risk assessment According to Bouch & Marshall, the following stages are identified for risk assessment and management. A. Identifying the need for a full structured risk assessment (not everyone will need this) B. Assessing static, stable, dynamic and future risk factors and considering protective factors C. Individual formulation of risk applied to the context of current presentation D. Considering possible interventions and the level of support required E. Anticipating the impact of possible interventions F. Developing a management plan with specified short and long term implementations G. Reviewing and revising the management plan with variations in risk factors. Commonly used tools Structured Risk Tools HCR-20 (Webster): It is a popular structured clinical assessment tool for violence risk. HCR-20 (Historical, Clinical and Risk) shows good inter-rater reliability. It has 10 historical items (history of previous violence, PCLR score, etc.) 5 Clinical items (lack of insight, diagnosis of PD) and 5 risk management items (feasibility of plans, lack of support, etc.) It has been useful in predicting inpatient violence and community violence in discharged patients. Historical items Clinical items Risk items Previous violence history Negative attitudes to health services Management plan lacks feasibility Young age at first incident Active symptoms Exposure to destabilisers (e.g. alcohol) Unstable relationships Impulsivity Non-compliance Major mental illness Treatment unresponsiveness Stress Substance use Lack of insight Lack of personal support Psychopathy
Employment issues Personality disorder Early maladjustment Previous supervision failure
19 - Actuarial instruments
Actuarial instruments
© SPMM Course SARA: Spousal assault risk assessment guide SARA is a 20 item set of risk factors for use in the assessment of spousal assault. It can be used to help gauge the risk of future violence in men arrested for spousal assault. SVR-20 is a sexual violence risk 20 scale - this is a 20 item guide for assessing violence risk in sex offenders. SAD PERSONS Score: 10 major demographic risk factors used in a mnemonic to assess immediate suicidal risk often in acute general hospital setting. The scores can guide in making a decision to admit or discharge a patient. S – Sex: 1 if male; 0 if female; (more females attempt, more males succeed) A – Age: 1 if < 20 or > 44 D – Depression: 1 if depression is present P – Previous attempt: 1 if present E –Ethanol abuse: 1 if present R – Rational thinking loss: 1 if present S – Social Supports Lacking: 1 if present O – Organized Plan: 1 if plan is made and lethal N – No Spouse: 1 if divorced, widowed, separated, or single S – Sickness: 1 if chronic, debilitating, and severe
Beck Hopelessness Scale consists of 20 true-false statements focused on pessimism and negativity about the future. The degree of hopelessness measured using this tool is a good indicator of suicidal risk with scores: 0 –3 indicating minimal, 4 – 8 mild, 9 –14 moderate, and 15– 20 severe risk. Beck Scale for Suicidal Ideation is a self-report 24-item scale (5 screening items) that assesses a patient’s thoughts, plans and intent to commit suicide. The total scores could range from 0 to 48 (each item scored from 0 to 2). Higher scores reflect greater suicide risk though no defined cutoffs are identified for categorizing the risk profiles. Actuarial instruments Group data is obtained from high-risk individuals and then to applied to the patient in question. It gives a group risk, and it should be applied with caution. Different types include: VRAG (violence risk appraisal guide – Quinsley 1995) entirely reliant on historical factors. Validated at Canadian prisons. It is made of 12 items and includes PCL_R as a subscale.
© SPMM Course Violence risk appraisal guide PCL-R Absence of schizophrenia (Presence is counted to decrease risk!) Elementary school difficulties Victim injury (minimal or none) Personality disorder Alcohol abuse Younger age Female victim Separated from parents before age 16 Failed conditional release/supervision order Never married History of non violent offence
Violence Risk Scale has 23 dynamic and 6 static variables. PCL-R (Hare) is a scale to diagnose Psychopathy, informs risk assessment and treatment decisions. 0 – 40 score range; 0-2 for each item; 20 items in total. Cut off of 25 used to diagnose psychopathy. In the strictest sense, PCL was not designed to be an actuarial tool for risk assessment on its own. The Static-99 is a ten item actuarial assessment instrument created by Hanson and Thornton, for use with adult male sexual offenders who are at least 18 year of age at the time of release to the community. SORA (Sexual risk offender appraisal guide) is a 14 item actuarial instrument that incorporates PCL_R. The Manchester Self Harm Rule (MSHR) is an actuarial instrument for self-harm risk assessment produced by Cooper et al. 2006. It has high sensitivity but low specificity.
© SPMM Course Notes produced using excerpts from: Achenbach, T. M. (1991). Manual for the Child Behavior Checklist/4-18 and 1991 Profile. Burlington, VT: University of Vermont, Department of Psychiatry. Bouch, J., & Marshall, J. J. (2005). Suicide risk: structured professional judgement. Advances in Psychiatric Treatment, 11(2), 84-91. Burns, A., Lawlor, B., & Craig, S. (2002). Rating scales in old age psychiatry. The British Journal of Psychiatry, 180(2), 161-167. Casey, P. & Kelly, B. (Ed) Fish’s Clinical Psychopathology. 3rd ed. RCPsych publications. Cooper J, Kapur N, Dunning J, et al. A clinical tool for assessing risk after self-harm. Ann Emerg Med. 2006;48:459–466. Cox JL, et al. Validation of the Edinburgh Postnatal Depression Scale (EPDS) in postnatal women. J Affect Disord 1996;39:185-9. http://www.static99.org/ http://www2.massgeneral.org/schoolpsychiatry/screeningtools_table.asp https://www.cnsforum.com/educationalresources/ratingscales/psychiatry Jackson. C. The General Health Questionnaire. Occupational Medicine 2007 57(1):79; Kaplan & Sadock's Synopsis of Psychiatry: Behavioral Sciences/Clinical Psychiatry, 10th Edition. Lippincott Williams & Wilkins 2007 Morgan et al.(1999), SCOFF Questionnaire. BMJ 319:1467 Sharp LK, Lipsky MS. Screening for depression across the lifespan: a review of measures for use in primary care settings. Am Fam Physician 2002;66: 1001-8. http://www.aafp.org/afp/2002/0915/p1001.html Ware Jr, J. E., & Sherbourne, C. D. (1992). The MOS 36-item short-form health survey (SF-36): I. Conceptual framework and item selection. Medical care, 473-483.
DISCLAIMER: This material is developed from various revision notes assembled while preparing for MRCPsych exams. The content is periodically updated with excerpts from various published sources including peer-reviewed journals, websites, patient information leaflets and books. These sources are cited and acknowledged wherever possible; due to the structure of this material, acknowledgements have not been possible for every passage/fact that is common knowledge in psychiatry. We do not check the accuracy of drug-related information using external sources; no part of these notes should be used as prescribing information.