NivaarExam PrepOfficial exam papers ↗

25-Comp-B10 Distributed Systems · May 2013

Question 6 of 7: Distributed File Systems

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

98-Comp-B10 Distributed Systems — National Examinations, May 2013. 3 hours, closed book, non-programmable calculator only. Candidates were instructed to answer any five of the seven questions, all carrying equal weight and mostly requiring essay-format answers; all seven are answered below as a complete study resource.

Reference texts: Coulouris, Dollimore, Kindberg & Blair, Distributed Systems: Concepts and Design (5th ed.) — system models and client-server architecture (ch. 2), interprocess communication and the request-reply protocol (ch. 4–5), operating system support for distributed systems (ch. 7), security (ch. 11), distributed file systems (ch. 12–12.4, AFS/NFS), and time, coordination, replication and fault tolerance (ch. 14–15, 18).

Check — sub-part lettering. Both sub-parts of Questions 1, 3 and 5 are lettered “a.” in the paper's numbering. Each of those three questions is answered below as two genuinely distinct sub-parts, relettered (a) and (b) in the order printed; content and marks weight are unaffected.

Question 6: Distributed File Systems

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

(a) Three key design issues for distributed file systems. 1. Naming and transparency. Clients should be able to name and access remote files using the same syntax as local files, ideally without knowing which server actually holds them (location transparency) and without needing to know if a file has moved (location independence). Getting this wrong forces applications to hard-code server identities, defeating the whole point of a shared, uniform file namespace. 2. Caching strategy and consistency. Because network access is orders of magnitude slower than local disk/memory access, clients must cache file data locally to be usable at all — but caching immediately raises the question of how (and how often) a client's cached copy is kept consistent with a file that another client may be concurrently modifying, trading off staleness against server/network load for every design choice made. 3. Fault tolerance and availability. A distributed file system must keep working (or degrade gracefully) in the presence of individual server or network failures, typically via some combination of replication (multiple copies of data on independent servers) and client-side resilience to a temporarily unreachable server, rather than an outage on one server making every client's work stop.

(b) The NFS Automounter. Rather than every client statically, eagerly mounting every remote filesystem it might ever need at boot time, the Automounter mounts a remote filesystem on demand, the first time a process actually references a path under it, and automatically unmounts it again after a period of inactivity. This improves performance because a client's boot/startup time and steady-state resource usage (open file handles, kept-alive connections) no longer scale with the total number of filesystems it might conceivably use, only with the ones actually in active use at any moment. It improves scalability in two ways: it avoids every client eagerly contacting every server it has ever been configured to reach (which would otherwise create an unnecessary storm of mount requests whenever a large fleet of clients reboots together), and it lets an administrator configure a client with access to a large, shared namespace of servers (e.g. via an indirect/auto map) without any per-client cost proportional to that namespace's size, since only the small subset actually touched at runtime is ever mounted.

(c) AFS vs. NFS — stability and scalability. Stability/consistency model: classic NFS (v2/v3) is essentially stateless and relies on clients periodically re-validating cached data against the server (a time-based, "check on open" or polling consistency), which means the server does no bookkeeping of who is caching what but pays a steady stream of validation requests from every active client, and different clients can transiently see different (stale) data between validations. AFS instead uses whole-file caching with server-issued callbacks: when a client caches a file, the server promises to notify (callback) that client if the file changes, so the client can trust its cache is valid until told otherwise, without needing to keep re-asking the server. This shifts work from "many clients constantly polling" to "the server does a small amount of bookkeeping and notifies rarely," which is inherently more stable and scalable under many, mostly-read clients, at the cost of the server needing to track callback state per client per cached file. AFS scalability limits: even with servers added as required, AFS's callback bookkeeping itself grows with the number of (client, cached-file) pairs a server must track, and a server recovering from a crash must re-establish (or invalidate) every outstanding callback it had promised, which becomes a heavier recovery burden as the client population grows; AFS's cell-based administrative partitioning also means very large deployments must be explicitly divided into cells, which is an organizational rather than a purely technical scaling limit. Recent developments: NFSv4 closed much of the original gap by adopting AFS-style delegations (a server-granted, revocable promise similar to a callback) plus a stronger, compound-RPC, more stateful protocol design, and it is now the more actively maintained/adopted standard; Coda (built directly on AFS's model) added disconnected operation, allowing clients to keep working from cache during a total server/network outage and reconcile changes afterward, addressing availability rather than raw throughput scalability.