Before I continue though, I want to state that I do believe in accountability and assessment, just not the way it is currently practiced. I believe strongly that good education can only come about when the Education Triangle is operating properly: that each of the parts act to inform the others in a dialectical manner. Because it is such a huge component of the current accountability movement, I want to focus on assessment today and in particular, how it can be used to inform teaching. Assessment is really about measurements. There are two ways you can think about using assessments:
- Measure what students already know.
- Measure what students have learned.
Assessment is not about standardized testing only, which to me is perhaps the least useful form of it. As a data driven educator, I depend on my own assessment data to gauge my growth as a teacher. For instance, take these two word clouds that I had my students generate about "science": the first one was at the beginning of a course on the scientific process (pre) and the second was from the end of the same course (post).
| Before and After Science Word Clouds from Spring 2013 |
The word clouds demonstrate the effect of instruction on my students as a whole group. What started off as a nebulous understanding of science, filled with generic science words, became refined and precise about the scientific process. It wasn't that students didn't have an understanding of science, but that it wasn't a clear one. By acknowledging what students' already know (prior knowledge), and using that to inform my lessons, I empowered my teaching to be more effective. As a bonus, notice that the word "boring" appeared in the pre, but "fun" appears instead in the post. That tickled me to no end. Other forms of non-testing based assessments can be things like a work portfolio demonstrating growth over the course of the class. My students have such portfolios that I ask them to review and reflect on at the conclusion of a course.
The use of a pre assessment to determine prior knowledge can often result in surprises. For instance, this was a question that I thought would be relatively hard for students since it tests for an understanding of inheritance mechanism and genetic recombination, but as it turns out, it was really easily and those few who didn't know, were able to learn it. "R" stands for right and "W" stands for wrong.
When asked about why this question was so easy, the consensus was that sex is needed to create babies and that children have the features of both parents. Knowing that this was how my students understood the question allowed me to better tailor the lesson to what they already know.
Pre-post assessments also allows for questions to be refined. Refining questions for future use is about understanding how they behave when deployed. Things to consider when revising a question are whether it tests for what we think and its difficulty level. The former is a bit complex so I'll save it for another time. There are two way to make a question more difficult:
The use of a pre assessment to determine prior knowledge can often result in surprises. For instance, this was a question that I thought would be relatively hard for students since it tests for an understanding of inheritance mechanism and genetic recombination, but as it turns out, it was really easily and those few who didn't know, were able to learn it. "R" stands for right and "W" stands for wrong.
| A surprising example of an easy question |
Pre-post assessments also allows for questions to be refined. Refining questions for future use is about understanding how they behave when deployed. Things to consider when revising a question are whether it tests for what we think and its difficulty level. The former is a bit complex so I'll save it for another time. There are two way to make a question more difficult:
- By asking students to use their knowledge in higher order thinking processes like application and analysis.
- By making the questions obtuse, confusing, or otherwise trick the student.
| A teach-able and learn-able question that is moderately difficult by requiring application of the Doppler Effect |
As you can see, even though many students got it wrong initially, a good number were able to learn how to apply the Doppler Effect to correctly answer the question. I could have made the question easier by asking for knowledge recall: "An object moving towards you would exhibit blueshift" but I wanted to activate higher ordered thinking skills to see if they could apply what they've learn. On the other hand, this next question is confusing and tricky.
| A difficult question because it tries to trick the student by using the word "translated" instead of "transcribed" |
Most students did not get the correct response either time, not because it was conceptually difficult, but because the question swapped two similar words. I believe that a better way to ask this question would be, "DNA is translated into RNA which is then transcribed into proteins." While still what I consider to be tricky, it at least makes salient how the words "transcribed" and "translated" are being used. The problem with making a question confusing or tricky is that it no longer becomes about testing for learning, but more about how proficient students are at figuring out what you are asking. What happens then is that you start measuring how well your students take tests rather than authentic learning. This defeats the point of having assessments to measure learning.
Some of you may be wondering why I used such a simplistic model to characterize questions than something more trendy like Item Response Theory (IRT). Aside from a distrust of complex models that requires to many assumptions to be made, I wanted to characterize individual items on their own as opposed to being a part of a set. IRT can only be used with a set of questions and the characterization is relative that set. Yes, I could do item equating, but I especially distrust that even more than I distrust IRT to be used properly. Futhermore, the questions are not testing for one concept or idea, and as such would violate the unidimensional assumption of IRT.
Some of you may be wondering why I used such a simplistic model to characterize questions than something more trendy like Item Response Theory (IRT). Aside from a distrust of complex models that requires to many assumptions to be made, I wanted to characterize individual items on their own as opposed to being a part of a set. IRT can only be used with a set of questions and the characterization is relative that set. Yes, I could do item equating, but I especially distrust that even more than I distrust IRT to be used properly. Futhermore, the questions are not testing for one concept or idea, and as such would violate the unidimensional assumption of IRT.
No comments:
Post a Comment