Selasa, 10 Maret 2015

STANDARDIZED PROFICIENCY TEST (PART II) : BULATS, TOEIC, TOEP By: I.G.A. Lokita Purnamika Utami Rina sari



STANDARDIZED PROFICIENCY TEST (PART II) : BULATS, TOEIC, TOEP
By: I.G.A. Lokita Purnamika Utami
Rina sari
This summary is the continuation of the previous summary about developing standardized test with elaboration on TOEFL and IELTS. However, this summary is now focusing on the other standardized test: BULATS, TOEIC, and TOEP which are related to job or business purposes. The last one can be also considered for academic purposes. 

Sabtu, 07 Maret 2015

Test wiseness: Definition, Types, and Implications, as well as Studies related with Test Wiseness

Test wiseness: Definition, Types, and Implications, as well as Studies related with Test Wiseness
By Rojab Siti R., & Muhammad Yunus

Gibb (1964) defined test-wiseness as the ability to respond advantageously to item clues in a multiple-choice setting and therefore to obtain credit without knowledge of the subject matter being tested. Test wiseness is also called test familiarity or test wisdom by Thorndike (1951: 569). It can lead to lower validity.
Types of test wiseness:

Selasa, 03 Maret 2015

Issues in the Development of Non-Test Instruments


Issues in the Development of Non-Test Instruments
 (Rating scales, semantic differential scales, checklists, questionnaires and others)

by: I.G.A. Lokita Purnamika Utami and Rina Sari

For the purpose of collecting new relevant data for a research study, the investigator needs to select proper instruments termed as tools and techniques. The major tools of research can be classified into broad categories of inquiry form, observation, interview, social measures and Psychological tests.
Among the inquiry forms, there are Rating scale, attitude scale, opinionnaire, questionnaire checklist and semantic differential scale.  Observation and Interview are explained as the techniques of data collection.  In psychological tests, Aptitude tests and inventories are discussed.

TEST WISENESS : DEFINITION, TYPES AND IMPLICATION AS WELL AS SOME STUDIES RELATED WITH TEST WISENESS



Marwa & Erlik Widiyani Styati
Defining Test-wiseness
Test-wiseness (TW hereafter) is a skill that permits a test-taker to utilize the characteristics and forms of tests and/ or test-taking situation to receive a high score. Some researchers (e.g., Benson, 1988; Rogers and Bateson, 1991) believe that TW is a cognitive ability or a set of test-taking strategies that a test taker can use to improve a test score no matter what the content area of a test. Bond (1981) distinguishes between test-wiseness and test-coaching. TW is independent of content areas whereas test-coaching refers to: “sustained instruction in the domain presumably being measured”.

Rabu, 25 Februari 2015

The summary of Developing standardized test of language proficiency: By: I.G.A. Lokita Purnamika Utami and Rina Sari



The summary of Developing standardized test of language proficiency:
By: I.G.A. Lokita Purnamika Utami and Rina Sari

Standardized test  for language proficiency presuppose a comprehensive definition of proficiency.Swain (1990) refers proficiency assessment to three linguistic traits: grammar, discourse and sociolinguistics that can be measure through oral, multiple choice and written responses. Another definition of proficiency is offered by ACTFL which offer a more holistic and unitary view: superior, advanced, intermediate and novice.

Sabtu, 21 Februari 2015

Quality Assurance on Internal Attributes of a Good Assessment Language Devices: Reliability, Validity, and Classical Item Analysis: Summary byMarwa & Erlik Widiyani Styati



Quality assurance on internal attributes of a good assessment consist of reliability, validity, and classical item analysis. Bachman (1990) says that reliability and validity are those two essentials to the interpretation and the use language ability. Besides, the is classical item analysis is also important. Here, the summary of quality assurance on internal attributes of good assesment will be described as follows:
1.      Reliability
Reliable means that the test can be trusted as a good test and it can be used many times and in the different time. Johnson and Johnson (2002) mention that reliability exists when students’ performance remain the same on repeated measurement. Reliability refers to the consistency of test scores; how consistent a particular students test scores are from one testing to another. Weir (1993) states that the test can be said to have high reliability if the result of the test shows the consistency when it’s re-used many times to a group of students in different time. The test can be said reliable if it is consistent. The types of reliability are: First, Inter-Rater or Inter-Observer Reliability which is used to assess the degree to which different raters/observers give consistent estimates of the same phenomenon. Second, Test-Retest Reliability is used to assess the consistency of a measure from one time to another. Third, Parallel-Forms Reliability is used to assess the consistency of the results of two tests constructed in the same way from the same content domain. Fourth, Internal Consistency Reliability is used to assess the consistency of results across items within a test. 

Selasa, 17 Februari 2015

A Summary on “Parallel Tests and Equating: Theories, Principles, and Practice”
By: Agus Eko Cahyono and Jumariati

In the context of language testing, parallel tests is an important issue. Multiple test forms are said to be parallel when they are as equal to one another as possible in terms of test specification like the type, form, content, purpose, and of statistical criteria like level of difficulty, discriminating power, and distracters. The common example is a school program which has two types of parallel tests: one is for the achievement test while the other is for those test-takers who need the retesting. In this case, the tests must be parallel as the function is the same that is to assess the achievement of the test-takers. High-stakes test like Ujian Nasional in Indonesian schools is used to be parallel with regards to the function to assess students’ learning achievement into certain level in spite of its administration that may be in different point in time throughout the country. Thus, the tests are different from one administration with another but the forms are still similar (equal). This implies that the assembly of multiple test forms should be designed very carefully and properly to ensure the fairness to each of the test-taker and at the same time to maintain the security of the tests.
In fact, there is still the possibility that the multiple test forms that have been developed are not similar; some differences in the statistical characteristics are still found. Therefore, equating methods to face this problem are needed. Equating parallel tests is an important issue in standardized tests to maintain the fairness among test-takers taking the test either at the same time or at different point in time. In order to be equal, Kolen and Brennan (2004) define several equality characteristics that need to be met. First, the equal construct requirement in which the tests to be equated must measure the same construct. If the tests’ constructs are different, they cannot be equated. Second is the equal reliability wherein the tests should yield reliable results. The third is the equal symmetry which means that the equating transformations must be symmetrical. Fourth, the equity requirement which deals with a matter of indifference to each test taker whether test form X or test form Y is administered. Finally is the population invariance requirement which means that the equating is the same regardless of the group of test-takers on which the equating is performed. These principles need to be taken into consideration once an equating is made.
In the practices of equating test forms, some methods are used such as the equating traditional method which utilizes the Item-Response Theory method and computer software like Kernell method and Automated Test Assembly (ATA). The traditional equating model is commonly done through random group design. In this type, test Model A is given to test-taker one, test Model B is given to test-taker two, and then test Model A is for test-taker three and so forth. The results obtained by the test-takers working on test Model A are compared to the result of those working on Model B. The conclusion then is made based on whether or not there is a difference between the two groups. If students in Model B obtain higher scores than those working on Model A, we can conclude that test Model B is easier than Model A and thus they are not equal (parallel).
The use of computer technology as ATA in assembling multiple test forms is preferred by test assemblers lately because of the fast processing and abundance of item pools (Lin, 2008). With the development of ATA, pre-equated parallel test forms can be achieved more efficiently.  The computer software will process the test criteria that have been laid out in the test blueprint and these criteria are separated into two: psychometric and non-psychometric attributes. Non-psychometric attributes include the test content, test format, test length, item usage frequency, and item exclusion. Meanwhile, psychometric attributes deal with classical item statistic, IRT-based item parameter estimates, item-response function, or item information functions.
In conclusion, equating multiple test forms is a crucial method of ensuring the equality of the tests; it can help test designers guarantee the fairness of the test to each of the test-taker and the security of the test forms.


References:
Kolen MJ, Brennan RJ.2004. Test Equating: Methods and Practices (2nd ed.). New York:
Springer-Verlag.

Lin, C.-J. 2008. Comparisons between Classical Test Theory and Item Response Theory
in Automated Assembly of Parallel Test Forms. Journal of Technology, Learning, and
Assessment, 6(8). Retrieved at February, 10th 2015 from http://www.jtla.org