25-Comp-B10 Distributed Systems · Undated paper
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.
(a) Transparency, fault tolerance and file-replication requirements for DFS design. Transparency requires that a file be named and accessed the same way regardless of which server physically stores it (location transparency), that moving a file between servers not require any client-side change (location independence), and that the interface for local and remote files be identical so existing single-machine applications work unmodified (access transparency) — together, these hide the fact that "the file system" is really many independent servers. Fault tolerance requires that the loss or temporary unavailability of one server not make files stored there permanently unreachable or, worse, silently corrupted; this is typically achieved by combining client-side caching (so brief server outages are masked entirely) with server replication and a well-defined recovery protocol (e.g. a stateless server design, so a crashed-and-restarted server has no fragile per-client session state to reconstruct). File-replication requirements go further, requiring multiple servers to hold up-to-date copies of the same file so that (i) reads can be served by whichever replica is closest/least loaded, improving both performance and availability, and (ii) the system defines a clear consistency model for what happens when a client writes to a replicated file — how quickly other replicas must reflect the update, and what a concurrent reader on another replica is guaranteed to see — since replication for availability directly trades against strict consistency unless carefully managed.
(b) Is NFS a distributed file system? Yes. NFS provides the defining properties of a distributed file system: it lets client machines mount and access files that physically reside on a remote server through the same file-system interface (open/read/write/close) used for local files, giving access and (largely) location transparency, and it allows many clients across a network to share the same files concurrently through one authoritative server. It is, however, a comparatively minimal/classic example of the category — it does not natively provide multi-server replication, automatic fault tolerance beyond retrying a stateless request, or the aggressive client-side whole-file caching with callback-based invalidation that a more elaborate DFS such as AFS provides — but minimal support for a property is not the same as absence of it, and NFS unambiguously satisfies the core definition of a distributed file system: multiple, independent client machines transparently sharing files stored on a server or servers elsewhere on the network.
(c) Stateful vs. stateless servers. A stateful server retains per-client session information across requests — e.g. which files a client currently has open, the client's current read/write offset within each file, and any locks the client holds — so a later request can be short (just "read the next N bytes") because the server already remembers the context. This makes ordinary operation more efficient (smaller messages, server can do read-ahead based on known access patterns) but makes crash recovery hard: if the server crashes, all of that per-client state is lost and must somehow be reconstructed (or every client's session broken) on restart. A stateless server, in contrast, keeps no memory of past requests at all: every request must be entirely self-describing (classic NFS, for example, has every RPC call carry the file handle, the exact byte offset, and the operation to perform). This makes each message larger and precludes some optimizations, but crash recovery becomes trivial — a restarted stateless server has no session state to lose, and clients simply retry their next self-contained request once it is back online.
(d) AFS vs. NFS — scalability, advantages and disadvantages. Scalability. AFS scales to substantially more clients per server than classic NFS because it has clients cache entire files locally on first open and relies on server-issued callbacks to notify a client only when a cached copy becomes stale (rather than the client re-checking with the server on every access); this dramatically reduces steady-state server load per client, which is why AFS deployments have historically supported many thousands of clients per server in large university/enterprise installations. Classic NFS instead validates cached data against the server far more frequently (traditionally via periodic polling/attribute checks on open or access), so per-client server load stays proportionally higher as the client population grows, limiting how far a single NFS server scales before another server (and its own separate namespace/mount point) must be added. Advantages/disadvantages. AFS's advantages are its scalability (per the callback mechanism above), a single, uniform, location-transparent global namespace across an entire organization (a client sees the same path everywhere, unlike NFS's per-server mount points), and strong support for disconnected/weakly-connected operation via its local whole-file cache; its disadvantages are the greater implementation/administrative complexity of maintaining server-side callback state (AFS servers are not fully stateless) and coarser consistency (whole-file caching means two clients writing the same file concurrently only see each other's changes on close, not continuously). NFS's advantages are its simplicity, its stateless-server design (trivial crash recovery, per part (c)), and its broad, mature cross-platform support as a long-established open standard; its disadvantages are weaker scalability under heavy concurrent load (per-client server checking cost) and a namespace that is only transparent within what an administrator has explicitly mounted, not organization-wide by default the way AFS's namespace is.