COMP 202 Foundations of Programming • McGill University, Montreal

Revision sheet: Files, exceptions and real data (COMP 202)

This sheet is not a summary of the file and exception lectures of COMP 202 at McGill University: you have the slides. It answers one question, what makes students lose marks on this chapter, on the paper midterm and final and in the assignments an autograder marks, and which precise gesture avoids each loss.

The angle of the chapter is that the world outside the program is text and it is never clean. Every loss of marks below is one of two decisions taken carelessly: converting too early, too late or not at all, and catching too much, too little or in the wrong place. Every wrong line is quoted as a student writes it, next to the line to write instead.

Mark this sheet as read or add it to your favourites: a free account, no password, keeps your read sheets and favourites from one visit to the next and tells you which chapter to tackle next. Create your space, an email is enough.

The thread of the chapter

Everything outside the program is text, and it is never clean: what comes in is a string until you convert it, what goes out is nothing until you turn it into characters, and every contact with the outside can fail. Two decisions decide the marks of the chapter: what to convert and when, and where to catch what.

This chapter is part of COMP 202, Foundations of Programming (McGill)

The essentials

Everything that comes in is a string, and it comes in with its newline

  • • open(name, 'r') returns a file OBJECT, not the text: a handle with a position, through which characters are read.
  • • for line in infile: hands the lines over one at a time, each one still ending with its newline, the two characters \n in a printed repr. 'Ana,8' read from a file is 'Ana,8\n', six characters, not five.
  • • The order inside the loop never changes: line = line.strip(), then the test for an empty line, then fields = line.split(','), then the conversion of each field you compute with, int or float, one by one.
  • • int('8') works and int('8\n') works too, because int tolerates surrounding whitespace; '8\n' == '8' is False and '8\n'.isdigit() is False. The comparison fails, not the conversion, which is why the bug survives a first test.
  • • The handle carries a POSITION. A loop that read to the end leaves it there: a second loop over the same handle sees nothing and raises nothing.
Ana,8\nBen,5\nposition at the open: the first loop reads 2 linesposition after that loop: a second loop reads 0 linesnothing moves it back except infile.seek(0) or a new open
One handle, one position: the first loop reads two lines and leaves the position at the end, so a second loop over the same handle reads nothing. Read into a list once, or seek(0).

Markers read the first line of the loop body: line = line.strip() before any test is the method mark. The autograder's hidden test is a file whose last field is compared, not converted, and it is the newline that fails it.

Everything that goes out is characters, and only what you write gets there

  • • Modes: 'r' reads, 'w' EMPTIES the file the moment it is opened, before any write, 'a' opens at the end and adds, 'x' refuses to overwrite an existing file.
  • • outfile.write(s) takes ONE str and adds nothing: no space, no newline, no conversion. outfile.write(8) raises TypeError. print(x, file=outfile) converts and adds the newline, like print on the screen.
  • • Characters written wait in a buffer; they reach the disk when the file is closed or flushed. with open(name, 'w') as outfile: closes on the way out of the block, exception included.
  • • The with closes the file, it does not catch anything: a FileNotFoundError raised by the open itself happens BEFORE the block is entered.
  • • One writer at a time: the open with 'w' goes outside the loop that writes, and a log that must survive between runs uses 'a'.

A program that never writes 'w' next to a file name that holds real data is a program that never destroys its own assignment data. Read with 'r', add with 'a', produce a report with 'w' opened once.

An exception is a value travelling up the calls, and catching it is a decision about WHERE

  • • try: holds the risky line and nothing else. except ValueError: runs only if that class was raised. else: runs only if nothing was raised. finally: runs in every case.
  • • except ValueError as err: gives the message in err; a bare except: catches KeyboardInterrupt and SystemExit as well, and hides the typos of the handler itself.
  • • Catch what the world does to you: a mistyped value, a malformed line, a missing file. Let a bug of your own, a TypeError or a ZeroDivisionError in code you wrote, climb to the top with its traceback.
  • • The layer that catches is the layer that can DO something: the input loop asks again, the reading loop skips and counts the line, the loop over files moves to the next file.
  • • raise ValueError('mark must be between 0 and 100, got ' + str(mark)) says which rule and which value: a function that cannot fix a problem raises it, it never returns a special value in its place.

On the design item of the final, the mark goes to the sentence that names the layer: 'parse_line cannot recover from a bad line, so it raises; read_file can skip it, so it catches'. Without that sentence a correct handler earns half.

The rules in table form

Each row reads from left to right: the assumptions, then the result. A red cell is not an answer, it is the finding that the form settles nothing and the instruction to rewrite it. Every case is followed by a worked example.

The exception names the broken contract: catch it, or let it climb

Read a line as: this line raises the class in the middle column, and the last column says whether a handler belongs there. A red line is a rule that does not exist: it is not an answer, it is the order to name the class.

The lineRaisesCatch, or let it climb?
int(text) on '7.5' ValueError catch at the input loop, ask again

Example: int('7.5') fails on the dot: float('7.5') gives 7.5 and int(7.5) gives 7, but int reads digits only.

fields[1] on ['Feb'] IndexError catch in the reading loop, count and skip

Example: 'Feb'.split(',') has 1 element, index 0 only; 'Feb,'.split(',') has 2, and fields[1] is the empty string ''.

open(name, 'r'), no such file FileNotFoundError catch where the name came from

Example: A name typed by the user is asked again; a name fixed in the code climbs, and the traceback names the file and the line, 1 open to fix.

'total: ' + 12 TypeError let it climb: a bug of yours

Example: 'total: ' + str(12) is 'total: 12', 9 characters; the version without str never runs, and no handler should hide that.

total / count, count is 0 ZeroDivisionError guard the count, never catch

Example: 0 records read: the mean does not exist, so return None or print 'no data', and never 0, which is a mean of something.

except: everything, Ctrl-C included safe because it catches all rule that does not exist

Example: In a loop that asks 100 times, the interrupt key is swallowed 100 times, and a NameError typed in the handler is reported as 'invalid input' for ever.

What to do: Name the class: except ValueError:, or at the very widest except Exception:, which lets KeyboardInterrupt and SystemExit through.

The test to apply before writing except: if the message would read 'I could not do this, so I skipped it', catch; if it would read 'something impossible happened', let it climb.

The mistakes that cost marks

These are the errors I correct most often in session. Each one costs marks on a paper, even when the reasoning behind it is right.

1. Opening a file with 'w' to look at it

the data of the whole assignment, and on paper the whole question

What not to write

“with open('marks.txt', 'w') as f: for line in f: print(line). I only read it, so the file is unchanged.”

What to write

“with open('marks.txt', 'r') as infile: for line in infile: print(line.strip()). Mode 'w' truncates at the open, before anything is written.”

Why: Truncation is a side effect of open, not of write: a file opened with 'w' and closed at once has 0 characters left. Reading from a handle opened with 'w' also raises, but by then the content is gone.

2. Reopening the output file inside the loop that writes it

the whole output item: the autograder finds one line where it expects three

What not to write

“for name in names: with open('report.txt', 'w') as outfile: outfile.write(name + '\n'). Each student is written, so the report is complete.”

What to write

“with open('report.txt', 'w') as outfile: for name in names: outfile.write(name + '\n'). One open, outside the loop, and the loop inside the block.”

passopen inside the loop1Ana2Ben3Chenpassopen outside the loop1Ana2Ana, Ben3Ana, Ben, Chenthe file after each passthe file after each pass
Open inside the loop, the file holds one name after every pass; open once outside it, the file grows to the three names. The mode is right in both, the position of the open is not.

Why: Every pass truncates what the previous pass wrote, so the finished file holds the last student only. Switching to 'a' inside the loop is the wrong repair: it also appends to whatever the previous RUN left.

3. Comparing a field that still carries its newline

2 marks on paper, and every hidden test that compares the last field

What not to write

“fields = line.split(','); if fields[1] == '8': count += 1. It never matches, so the file must be corrupt.”

What to write

“line = line.strip() as the first line of the loop, then fields = line.split(','): fields[1] is now '8' and the comparison is True.”

Why: The last field of a line read from a file is '8\n', six characters split into two; int('8\n') works, which is why the bug is invisible until a comparison or an isdigit test is written.

4. Looping over the same handle twice

the whole computation item, with a total of 0 that looks like a conversion bug

What not to write

“for line in infile: count += 1, then for line in infile: total += int(line). The count is 3 and the total is 0, so int must be failing silently.”

What to write

“lines = [line.strip() for line in infile if line.strip() != '']; then count = len(lines) and the total is computed over lines, as many times as needed.”

Why: The first loop leaves the position at the end of the file; the second starts there and finds nothing, and nothing is raised. Reading into a list once is the answer; infile.seek(0) works and is rarely worth it.

5. Handing write a number, or forgetting the newline

1 mark for the TypeError, the whole file item for the glued lines

What not to write

“outfile.write(name + ',' + avg), then on the next line outfile.write(name + ',' + str(avg)). The first crashes, the second gives Ana,7.5Ben,6.0Chen,8.0 on one line.”

What to write

“outfile.write(name + ',' + str(avg) + '\n'), or print(name, avg, sep=',', file=outfile), which converts and ends the line.”

Why: write takes one str and adds nothing at all: a file is a stream of characters, and how a number should look in it is the program's decision. Three writes of 'Ana,7.5' give 21 characters and 0 newlines.

6. Writing a bare except because it is 'safer'

2 style marks on every assignment, and an input loop the user cannot interrupt

What not to write

“try: value = int(text) except: print('invalid input'). It catches everything, so the program can never crash.”

What to write

“try: value = int(text) except ValueError: print('Please type a whole number'). Name the class you expect, and nothing else.”

Why: A bare except catches KeyboardInterrupt and SystemExit, so the loop can no longer be stopped, and it catches the NameError of a typo in the handler, reported as invalid input for ever. except Exception: is the widest honest form.

7. Putting the whole body of the function inside the try

half the design item, and a real bug reported as bad input

What not to write

“try: value = int(text); result = compute(value); print(report(result)) except ValueError: print('not a number').”

What to write

“try: value = int(text) except ValueError: continue else: result = compute(value). The try holds the risky line; the rest goes in the else or after the block.”

Why: A ValueError raised much later by compute, for a completely different reason, is caught by this handler and announced as a bad number. The else clause exists precisely to keep the try short.

8. Turning an empty field into zero

the whole result item: a mean of 33 instead of 44 on the example file

What not to write

“if fields[1] == '': mm = 0 else: mm = int(fields[1]). The file has a hole, so I fill it with 0 and the loop goes on.”

What to write

“except ValueError: skipped += 1; continue. A reading that was never taken is not 0, it is absent: skip the record and count it, or store None.”

Why: A zero is a measurement that never happened, and it enters every sum and every mean. On rain.csv with 42, an empty field, 57 and 33, the mean over the three real readings is 44; with a zero for February it becomes 33.

9. Catching your own bug so that the program does not crash

the whole design item on the final, where the question is exactly this one

What not to write

“try: mean = total / count except ZeroDivisionError: mean = 0. The program must not crash, so I catch it.”

What to write

“if count == 0: return None, else return total / count. A ZeroDivisionError in code you wrote is a bug or an empty input, and both deserve a decision, not a handler.”

Why: Catching an exception that means the program's own assumptions are broken lets it carry on with a state it does not understand and print a wrong number. A guard on the count is a decision; a handler that sets 0 is a lie.

Which method to choose

Where the handler goes, by who can do something about it

Look at where the bad value came from, and what the code at that layer can do next

main(): loops over the file namesread_file(name): loops over the linesparse_line(line): int(fields[1])ValueError: caught in read_file, line counted and skippedFileNotFoundError: caught in main, message, next fileZeroDivisionError: caught nowhere, traceback, fix the code
Three exceptions, three answers: the ValueError stops one frame up, in the loop that can skip the line; the FileNotFoundError in the loop that can move on; the ZeroDivisionError stops nowhere.
  • If the value was typed by the user: int(input(...)) fails → catch ValueError in the input loop itself, print what is expected, ask again

    Example: while True: try: value = int(text) except ValueError: continue; then the range test OUTSIDE the try

  • If the line came from a file and does not parse: int on a field fails → let the parsing function raise; catch in the loop over the lines, count the line, continue

    Example: except ValueError as err: problems.append(number); continue, and 1 line skipped out of 5 is reported at the end

  • If the file itself is missing: open raises FileNotFoundError → catch in the loop over the file names, print a message naming the file, move to the next file

    Example: except FileNotFoundError: print('missing:', name); with 3 names and 1 missing, 2 files are read

    if the name is fixed in the code, do not catch: the traceback already names the file

  • If the error is in your own arithmetic or types: 'total: ' + 12, total / 0 → catch nowhere; guard the count if an empty input is possible, and fix the code otherwise

    Example: if count == 0: return None; the ZeroDivisionError never happens, and no handler hides a real bug

If no branch applies, the exception is not an event of the outside world, and the answer is no handler at all. A try that appears at every layer is the sign that the layer was chosen nowhere.

Which reading form, by what the program needs from the file

Look at what the rest of the program will do with the content

  • If one record per line, processed once, in order → for line in infile: with strip on the first line of the body

    Example: 4 data lines give 4 passes, each on a string ending with its newline until the strip

  • If the lines must be walked twice, or sorted, or counted before use → lines = [line.strip() for line in infile if line.strip() != ''], then work on the list

    Example: len(lines) is 4 and the same list feeds the total, the maximum and the report

  • If the first line is a header → infile.readline() once, before the loop, and nothing done with its value

    Example: 'month,mm' is consumed by readline, and the loop starts at 'Jan,42'

  • If the whole content is searched as one text → text = infile.read(), then the string methods of chapter 3

    Example: text.count('\n') is 4 on a file of 4 lines ending with a newline, 3 if the last line has none

    readlines() is rarely the best of the three: it holds every line with its newline

Whatever the form, the file is closed by the with, and the structure built from it, the list of records, is what the rest of the program uses. Nothing is computed while the file is still open except the reading itself.

How the answer is expected to be written

A marker ticks steps. Here they are in order, with the concluding sentence expected word for word.

A reading function that survives its user

When to use it: The statement asks for a function that reads an integer between two bounds and keeps asking until it gets one

  1. 1 Open a while True loop, and read the text with input(prompt): the loop is what makes the function survive, not the try.
  2. 2 Convert inside a try that holds ONLY the conversion: value = int(text), then except ValueError: print a message that names what is expected, and continue.
  3. 3 Test the range OUTSIDE the try, after the except: if lo <= value <= hi: return value, else print the bounds and let the loop run again.
  4. 4 State in one sentence why the range test is outside: a value out of range is not a conversion error, and the handler must not be able to catch it.

Concluding sentence

“The try holds int(text) only: a ValueError means the text was not a whole number and the loop asks again; the range test comes after the except, so that an out of range value is refused by the test and never by the handler.”

The trap: except: instead of except ValueError:, which turns the loop into one the user cannot interrupt.

Marking: Typically 1 mark for the loop, 1 for the try around the single line, 1 for the named class, 1 for the range test outside, 1 for the message.

From a CSV file to a number: read, clean, convert, then compute

When to use it: The statement gives a file of records, one per line with a separator, and asks for a total, a mean, a maximum or a report

  1. 1 Open with 'r' in a with, and consume the header with one infile.readline() before the loop if the file has one.
  2. 2 In the loop: line = line.strip() first, then if line == '': continue, then fields = line.split(',').
  3. 3 Convert the fields you compute with, one by one, inside a try that holds the conversions only; except ValueError: count the line and continue.
  4. 4 Append the converted record to a list; leave the with block, and compute totals, means and the report from the LIST, with the count guarded before any division.
  5. 5 Write the report with print(..., file=outfile) or with an explicit newline at the end of every write, the open with 'w' placed once, outside the loop.

Concluding sentence

“Each line is stripped, split on the comma and converted field by field; a line that does not convert is counted and skipped; the mean is computed from the list of valid records after the file is closed, and only if the count is not zero.”

The trap: Computing the mean inside the with block from a running total, then reading the file again for the report: the second loop over the handle sees nothing.

Marking: Typically 1 mark for the header, 1 for strip and the empty line, 1 for the conversion in a short try, 1 for the guarded mean, 1 for the newline in the output.

Check before you hand in

Five minutes of checking recover more marks than one more problem started in a hurry.

The typical problem, taken apart

A mean from rain.csv, with a header, an empty field and no final newline

The file rain.csv holds a header line, then one line per month in the form month,mm. Write mean_rain(filename), which returns the mean of the readings it can use, prints one message per line it cannot, and returns None when there is no usable reading. Do not catch the case of a missing file inside the function, and say why.

The file as it sits on disk: month,mm then Jan,42 then Feb, then Mar,57 then Apr,33, each line ending with a newline except the last.

month,mm\nheader line: readline() once, before the loopJan,42\nFeb,\nempty field: int('') raises ValueErrorMar,57\nApr,33no final newline: strip() makes it harmlessone run of characters: the newline is a character in the line
Five lines, three traps: the header is not data, the February line has an empty second field, and the last line has no newline. Only the second one needs a handler.
python
def mean_rain(filename):
    total = 0
    count = 0
    number = 1
    with open(filename, 'r') as infile:
        infile.readline()
        for line in infile:
            number += 1
            line = line.strip()
            if line == '':
                continue
            fields = line.split(',')
            try:
                mm = int(fields[1])
            except ValueError:
                print('line', number, 'skipped:', line)
                continue
            total += mm
            count += 1
    if count == 0:
        return None
    return total / count

Step 1

Open with 'r' inside a with, and call infile.readline() once before the loop, without keeping its value: the header 'month,mm' is consumed and the loop starts on 'Jan,42'.

Why

The header would otherwise reach int('mm') and be reported as a bad line, which it is not: it is structure, and structure is skipped on purpose, not caught by accident. The with guarantees the close, exception or not.

Step 2

First line of the loop body: line = line.strip(). Then if line == '': continue, then fields = line.split(','). The last line, 'Apr,33' with no newline, and the others, with one, come out identical.

Why

Strip before anything else is the one gesture that makes the newline, the missing newline and a Windows carriage return all disappear. The empty line test comes after the strip, since a line of spaces is empty too.

Step 3

The conversion mm = int(fields[1]) is the only line inside the try. On 'Feb,' the fields are ['Feb', ''], int('') raises ValueError, the handler prints line 3 skipped: Feb, and continues.

Why

The handler is here, in the loop over the lines, because this is the layer that can do something useful: skip the record and move on. Nothing else is inside the try, so a ValueError from elsewhere could never be reported as a bad line.

Step 4

total and count are updated after the try, on the valid lines only: 42, 57 and 33 give total 132 and count 3. After the with block, if count == 0: return None, else return total / count, which is 44.0.

Why

The guard on the count is a decision written in the code, not a ZeroDivisionError caught after the fact; and returning None for an empty file is honest where 0 would be a mean of something. The division is outside the with: the file is closed first.

Step 5

Check: with a zero stored for February instead of a skip, the mean would be (42 + 0 + 57 + 33) / 4 = 33 rather than 44; with the header not skipped, the first message would name line 1, 'month,mm'.

Why

Naming what each shortcut would have produced is the fastest way to be sure none was taken: a wrong mean of 33 is silent, the spurious message on line 1 is visible in the output.

Step 6

The missing file is not caught here. open raises FileNotFoundError before the with block is entered, and the exception climbs to the caller, who knows where the name came from and can ask again or move to the next file.

Why

mean_rain cannot resolve a missing file: it has no other name to try. A handler that printed a message and returned None would hide the difference between a missing file and an empty one.

The conclusion, written out

“mean_rain skips the header with one readline, strips every line before testing it, converts the reading inside a try that holds the int only, counts and skips the February line, and returns 132 / 3 = 44.0 after the file is closed; a missing file raises to the caller, who is the one able to act.”

The classic mistake on this problem: Writing mm = int(fields[1]) with the whole loop body inside the try, then except: print('bad line'). The header, the empty field and any later bug all print the same message, and the interrupt key no longer stops the program.

Learn by heart

  • • open(name, 'r') returns a HANDLE with a position; 'w' EMPTIES the file at the open; 'a' adds at the end.
  • • A line read from a file ends with its newline: line = line.strip() is the first line of the loop body, before any test or split.
  • • A handle read to the end gives nothing on a second loop and raises nothing: read into a list once.
  • • write takes one str and adds NOTHING; print(x, file=outfile) converts and ends the line. The open with 'w' goes outside the loop.
  • • try holds the risky line only; except names the class; else runs when nothing was raised; finally runs always. Never a bare except.
  • • Catch what the world does, at the layer that can act: ask again, skip the line, move to the next file. Let your own bugs climb.
  • • An absent value is not 0: skip and count it, or store None. Guard the count before dividing.

Frequently asked questions

Why does my comparison fail on the last field of a line read from a file in Python?

Because the line still ends with its newline character. When you split 'Ana,8' followed by a newline on the comma, the last field is '8' plus the newline, so comparing it to '8' gives False even though converting it with int works, since int ignores surrounding whitespace. Call line.strip() on the first line of the loop body, before the split, and every field comes out clean.

Why is my file empty after opening it with mode w in Python?

Mode w truncates the file to zero length the moment open is called, before anything is written, so a file opened with w to look at it is emptied even if the program never writes and even if it crashes right after. Use r to read, a to add at the end of an existing file, and w only for a report you produce from scratch, opened once outside the loop that writes it.

Why does my second loop over a file in Python read nothing?

A file object keeps a position, and the first loop moved it to the end of the file. The second loop starts there, finds no line, and raises nothing, so the symptom is an empty result rather than an error. Read the file into a list of stripped lines once, inside the with block, and loop over that list as often as you need; infile.seek(0) also works but is rarely the clearer answer.

Should I use a bare except in a COMP 202 assignment?

No. A bare except catches every exception, including the one raised by the interrupt key and the one raised when the program asks to exit, so a loop that asks the user again can no longer be stopped. It also catches a typo inside the handler itself and reports it as invalid input. Name the class you expect, usually ValueError; if you really need a wide net, except Exception is the widest honest form.

Where should I catch an exception in a program that reads several files?

At the layer that can do something useful with it. A line that does not convert is caught in the loop over the lines, which can count it and skip it. A missing file is caught in the loop over the file names, which can print a message and move to the next one. A function that cannot fix the problem, like the one parsing a single line, raises instead of catching. A bug in your own arithmetic is caught nowhere: let the traceback reach you and fix the code.

Practise it

Corrected exercises: Files, exceptions and real data, COMP 202

A method is proved on a paper, not on a sheet. The set for the same chapter takes each of these traps into a problem, with the solution written out step by step.

  • 10 corrected exercises
  • 100 points
  • 150 minutes
Do the exercises
Previous sheet Lists, dictionaries and the structure to choose Next sheet Recursion, and when a loop is the better answer

© Ahmed Squalli Houssaini. Revision sheet published at www.letuteurscientifique.ca/en/fiches/comp202-files-and-exceptions. Free for personal and classroom use; republishing it elsewhere requires written permission (legal notice).

See also

Stuck in COMP 202?

I tutor COMP 202 at McGill in English or in French, in person in Montreal or online. Get in touch for a first session: the file and exception assignment is the one where a working program can still lose most of its marks.

Site by Studio Squalli