LƯỢT TRUY CẬP

  • truy cập   (chi tiết)
    trong hôm nay
  • lượt xem
    trong hôm nay
  • thành viên
  • CHUYÊN MỤC CHÍNH

    Test Usefulness

    Wait
    • Begin_button
    • Prev_button
    • Play_button
    • Stop_button
    • Next_button
    • End_button
    • 0 / 0
    • Loading_status
    Nhấn vào đây để tải về
    Báo tài liệu có sai sót
    Nhắn tin cho tác giả
    (Tài liệu chưa được thẩm định)
    Nguồn: internet
    Người gửi: Vũ Trung Kiên (trang riêng)
    Ngày gửi: 15h:21' 10-09-2014
    Dung lượng: 657.0 KB
    Số lượt tải: 6
    Số lượt thích: 3 người (Nông Đức Hạnh, Ông Cao Thắng, Ngô Thừa Ân)
    1




    Module 5119
    Language Testing and Assessment

    Week 2
    Test usefulness model
    2
    Points to be covered

    Brief look at Week 1
    Model of test usefulness
    Preparation for Discussion Leader # 1
    From Week 1
    Definition of test, its uses and purposes

    Types of tests: pen-and-paper tests and performance tests; achievement tests and proficiency tests

    Test-criterion relationship




    4






    Based on the obvious - both usefulness and usability.

    Provides a framework for evaluating the whole process of test development and use and ensuring quality control.

    Considers the question of ‘fairness’ as well as social sensitivity.

    Emphasizes testing as a contextualized and situated practice.




    5
    Bachman and Palmer’s (1996)
    model of test usefulness
    The model
    Usefulness = Reliability + Construct validity +
    Authenticity + Interactiveness +
    Impact + Practicality

    Operationalizing principles
    1 Maximizing overall usefulness, rather than individual test qualities
    2 Interdependence of the qualities
    3 Appropriate balance is context-dependent
    6
    Reliability vs. Validity
    Reliability refers to the consistency of the results obtained from a piece of research.
    Validity has to do with the extent to which a piece of research actually investigates what the researcher purports to investigate.
    Nunan (1992, p. 14)
    Reliability? Can we trust the measurements to be accurate?)
    Validity? Have we measured the abilities that we set out to measure? (e.g., reading comprehension ability, level of student interest in a learning task, or writing ability)
    7
    A reliable test?
    Is consistent and dependable: give the same test to the same student or matched students on two different occasions, the test yield similar results
    8
    Factors contributing to the unreliability of a Test
    Student-related reliability
    e.g., Temporary illness, physical or psychological factors
    Rater reliability
    e.g., subjectivity, bias
    Inter-rater reliability: inconsistent scores
    Intra-rater reliability (not even-handed judgment): faced with up to 40 tests to grade in only 1 week
    Test administration reliability
    The conditions in which the test is administered
    Test reliability
    Test is too long
    Test items ambiguous, or have more than one correct answer
    9
    Reliability
    Consistency of estimation of candidates’ abilities

    Consistency of measurement across different occasions, settings and versions of a test.

    Quality of test scores. Unreliable scores do not constitute sound evidence for making inference about test takers’ language ability; cannot draw any conclusions

    A necessary element of/condition for construct validity and thus test usefulness.

    Inconsistencies cannot be eliminated entirely; measurement errors should be minimized.
    10
    Validity
    “The extent to which inferences made from assessment results are appropriate, meaningful, and useful in terms of the purpose of the assessment” (Gronlund, 1998, p. 226) (cited in Brown, 2004)
    “Test validity is predicated on there being a lack of bias in tasks, items, or test content” (Ross & Okabe, 2006. p. 1)
    Gronlund, N. E. (1998). Assessment of student achievement (6th edn). Boston: Allyn and Bacon.
    A valid language test?
    measures accurately what it is intended to measure
    e.g., a valid test of reading ability actually measures reading ability not writing ability
    Reflection
    to measure writing ability
    T asks S to write as many words as they can in 15 minutes, then simply counts the words for final score → easy to administer (practical) and the scoring quite dependable (reliable)
    → not a valid test: cos’ without consideration of comprehensibility, rhetorical discourse elements, organization of ideas, etc.
    13
    How is the validity of a test established?
    There is no final or absolute measure of validity
    However, validity is established:
    It examines the extent to which a test calls for performance that matches that of the course or unit of study being tested.
    Determines whether or not S have reached an established set of goals or level of competence
    Statistical correlation with other related but independent measures
    Aspects of Validity
    Validity is a complex concept and reveals a number of aspects:
    Construct validity
    Face validity
    Content validity
    Criterion-related validity
    Consequential validity
    Construct Validity
    Construct: An abstract concept that has theoretical existence but cannot be observed directly (e.g., intelligence, communicative competence, motivation, self-esteem)

    (construct = a trait or underlying abilities)

    Meaningfulness and appropriateness of the interpretation that is made on the basis of test scores.

    Justify the meaning of scores by means of evidence and logic.
    16
    IELTS scores interpretation
    17
    Construct validity (Cont.)
    Language knowledge is defined as constructs (e.g. linguistic competence; communicative competence, proficiency)
    e.g., ability to read involves ability to guess the meaning of unknown words from the context
    Domain-specific construct (e.g. ability to write for a discourse community; ability to communicate in specific workplaces)
    Definition of construct is based on theories of language knowledge and competence.
    e.g., theory of writing tells us that underlying writing ability involves control of punctuation, sensibility to demands on style ; or
    Oral proficiency: pronunciation, fluency, grammatical accuracy, vocab. use, social linguistic appropriateness
    Compatibility of defined construct and target language use (TLU) tasks
    Correspondence between test tasks and TLU tasks: The construct measured by test tasks should reflect the construct defined, or underlying TLU tasks.
    Appropriate representation of construct in test tasks; avoidance of construct under-representation or construct-irrelevance.
    Performance measures (i.e. test scores) are indicators of corresponding degrees of language competence or task performance.
    Generalizing test performance to TLU domain.





    18
    Reflection Point
    Case 1: You are asked to conduct an oral proficiency interview that evaluates only pronunciation and grammar.
    Construct validity of the test?
    Case 2: You have created a simple written vocab. quiz, covering the content of a recent unit, that asks sts to correctly define a set of words. The lexical objective of the unit is communicative use of vocab.
    What are your views on your chosen items?
    Construct Validity (Cont.)
    Test validation is an on-going process, and interpretations of scores can never be absolutely valid.

    Unlike reliability, which can be tested by statistical measures, validity is logical, judgmental and empirical.


    20
    Face validity (FV)
    FV refers to the degree to which a test looks right, and appears to measure what it is supposed to measure
    e.g., a test to measure pronunciation ability, but not require S to speak
    This is true for the test’s construct and criterion-related validity
    Face validity is hardly a scientific concept/empirically tested, yet it is important: a test without FV may not be accepted by candidates, teachers, education authorities or employers.
    Face Validity (cont.)
    FV will likely be high if learners encounter:
    A well-constructed, expected format with familiar tasks,
    A test that is clearly doable within the allotted time limit,
    Items that are clear and uncomplicated,
    Directions that are crystal clear,
    Tasks that relate to their course work (content validity),
    A difficulty level that presents a reasonable challenge
    Content validity (CV)
    A test’s content constitutes a representative sample of the language skills, structures, etc with which it is meant to be concerned (Mousavi, 2002; Hughes, 2003)
    E.g., If trying to assess S’s ability to speak E. in a conversational setting → ask S to answer paper-and-pencil multiple choice requiring grammatical judgments → not achieve CV
    Or, A course has 10 objectives, but only two covered in a test → not achieve CV
    Discussion Activity
    Consider the following quiz on E. articles for a high-level beginner of a conversation class (L & S) for E. learners
    S had had a unit on zoo animals and had engaged in some open discussions and group work in which they had practiced articles, all in listening and speaking modes of performance
    English articles Quiz
    Comments
    The quiz uses familiar setting and focuses on previously practiced lang. forms somewhat content validity
    The fact that it was administered in written form, and required S to read the passage and write their responses → makes it quite low in content validity for Listening and Speaking class
    Criterion-related validity
    The extent to which the ‘criterion’ of the test has actually been reached
    See how far results on the test agree with those provided by some independent and highly dependable assessment of the candidate’s ability.
    i.e, the results of one teacher’s unit test might be compared with possibly a commercially produced test in a textbook
    Or,
    scores of a test on gram. in communicative use are corroborated either by (a) subsequent behaviour or (b) other communicative measures of the grammar.
    Criterion-related validity falls into one of two categories:
    Consequential validity
    Considerations:
    A test’s accuracy in measuring intended criteria
    Its impact on the preparation of test-takers
    (e.g., socioeconomic conditions: opportunities for coaching; get help from educated parents)
    Its effect on learners
    Intended and unintended social consequences of a test’s interpretation and use
    Washback: effect of assessments on students’ motivation, subsequent performance in a course, independent learning, study habits, attitude toward school work
    Authenticity
    “The degree of correspondence of the characteristics of a given language test task to the features of a TLU task” (Bachman & Palmer, 1996, p.23)
    (One task is likely to be enacted in the ‘real world’)
    Relatively underemphasized in language testing, but recently increased noticeably

    Test takers’ perception of authenticity mediates their response to test tasks.

    Authenticity is essential for construct validity.

    Authenticity in relation to TLU domain or the language instruction domain (i.e. syllabus)?
    30
    Authenticity (cont.)
    Authenticity may be present in the following ways:
    The language in the test is as natural as possible.
    Items are contextualized rather than isolated.
    Topics are meaningful (relevant, interesting) for the learner.
    Some thematic organization to items is provided such as through a story line or episode
    Tasks represent, or closely approximate, real-world tasks.
    Interactiveness
    Whether and to what extent test tasks call for test takers’ language knowledge and ability for successful task completion.

    Interaction between test takers’ language ability (language knowledge, strategic competence, or metacognitive strategies), topical knowledge and affective schemata and task characteristics.

    Eliciting language performance data, and nothing else.

    Essential for construct validity.
    32
    Authenticity and Interactiveness

    Screening tests for typists which require copying from handwritten document.

    Screening tests for typists based on oral interviews about mundane topics.

    Vocabulary matching task for testing ability to read
    academic texts given to prospective students entering an American university.

    Role-play task about selling products for prospective salespersons.



    33
    Authenticity and Interactiveness (Cont.)
    Both authenticity and interactiveness are inherent in construct validity.

    They are relative — relatively more or relatively less authentic or interactive.

    Contingent on characteristics of test takers, test tasks and TLU tasks.

    Acceptable levels of authenticity and interactiveness depend on specific testing situation.






    34
    Impact
    Test impact refers to the social dimension of language testing and is also known as consequential validity.

    Test administration and use involves values and goals, the consequences of which affect individuals, education systems as well as the society at large.

    Test impact operates at both micro level and macro level.

    35
    Impact (Cont.)
    Micro level:
    Impact of particular test use on individuals
    (e.g. test takers, test users, decision makers, parents, educators, employers)

    Macro Level:
    Impact on society and educational systems

    36
    Impact (Cont.)
    Impact on test takers
    Test preparation and test taking experience
    e.g., A writing test only sets two kinds of tasks: compare/contrast; describe/interpret charts → much preparation for the test limited to those tasks – not beneficial backwash
    Feedback from tests
    Test results: inclusion, exclusion
    Washback
    ‘Teaching to the test’; opportunity for instructional reforms
    Impact on teachers
    Career; instructional preferences, etc.



    37
    Practicality
    Practicality refers to practical issues of test development and implementation.

    Consideration of the resources which determines the nature, type and length of tests.

    Relationship between the resources that are required in the design, development and use of the test and the resources that will be available for this purpose.

    Three types of resources are required:
    Human resources
    Material resources
    Time

    38
    Practicality (cont.)
    How long it will take to sit and to mark the test
    stay within appropriate time constraints;
    scoring/evaluation procedure is specific and time-efficient
    Physical constraints of the test situation: e.g., speaking test
    A test is of practicality: Fomula
    Ar: available resources
    Nr: needed resources
    Quotient ≥ 1, a test is of practicality; ≤ 1 not practical
    (Bachman & Palmer, 1996)
    A test should be easy and cheap to construct, administer, score and interpret
    39
    Summary-test usefulness
    Test usefulness incorporates all essential test qualities to ensure fairness and quality control.

    Social dimension of language test is emphasized in the model by including impact as an important test quality.

    Test usefulness ultimately depends on the purpose and the context of the test which also aims at a balance among the test qualities.

    40
    Practice 1
    Consider the following two excerpts from tests and evaluate them on a measure of authenticity
    Comments
    1st excerpt (contextualized tasks)
    The sequence of items achieves a modicum of authenticity by contextualizing all the items in a story line.
    might occur in the real world, even if with a little less formality
    2nd excerpt (decontextualized tasks)
    Sequence of items takes the test takers into five different topic areas with no context for any. Each sentence likely written or spoken in the real world, but not in that sequence
    The first is “good”; second is “fair”
     
    Gửi ý kiến

    XIN CHÀO VIỆT NAM

    CHÀO XUÂN BÍNH NGỌ 2026

    Chào mừng quý vị đến với CHÚT LƯU LẠI.

    Quý vị chưa đăng nhập hoặc chưa đăng ký làm thành viên, vì vậy chưa thể tải được các tư liệu của Thư viện về máy tính của mình.
    Nếu đã đăng ký rồi, quý vị có thể đăng nhập ở ngay ô bên phải.

    Nhúng mã HTML

    Nhúng mã HTML