NivaarExam PrepOfficial exam papers ↗

25-Comp-B10 Distributed Systems · May 2017

Question 5 of 6: Distributed File Systems

Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)

Notes on this paper

98-Comp-B10 Distributed Systems — National Examinations, May 2017. 3 hours, closed book, non-programmable calculator only. Candidates were instructed to answer any five of the six questions (only the first five as they appear in the answer book are marked), all carrying equal weight and mostly requiring essay-format answers; all six are answered below as a complete study resource.

Reference texts: Coulouris, Dollimore, Kindberg & Blair, Distributed Systems: Concepts and Design (5th ed.) — system models, peer-to-peer systems, middleware and client-server architecture (ch. 1–2), interprocess communication and the request-reply protocol (ch. 4–5), remote invocation (ch. 5), operating system support for distributed systems (ch. 7), security (ch. 11), distributed file systems (ch. 12), and time, coordination, replication and fault tolerance (ch. 14–15, 18).

Check — source parsing artifact. Every question header on this paper is printed as “Question # N.” (a literal hash between the word and the number).

Question 5: Distributed File Systems (20 marks)

Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.

(a) Three requirements in the design of distributed file systems. 1. Naming and transparency. Clients should be able to name and access remote files using the same syntax as local files, ideally without knowing which server actually holds them (location transparency) and without needing to know if a file has moved (location independence). Getting this wrong forces applications to hard-code server identities, defeating the whole point of a shared, uniform file namespace. 2. Caching strategy and consistency. Because network access is orders of magnitude slower than local disk/memory access, clients must cache file data locally to be usable at all — but caching immediately raises the question of how (and how often) a client's cached copy is kept consistent with a file that another client may be concurrently modifying, trading off staleness against server/network load for every design choice made. 3. Fault tolerance and availability. A distributed file system must keep working (or degrade gracefully) in the presence of individual server or network failures, typically via some combination of replication (multiple copies of data on independent servers) and client-side resilience to a temporarily unreachable server, rather than an outage on one server making every client's work stop.

(b) Which supports more clients on identical hardware: NFS or AFS. AFS supports more clients on a server of identical hardware. Classic NFS is essentially stateless and relies on clients periodically re-validating cached data against the server (time-based, "check on open"/polling consistency), so the server does no per-client bookkeeping but pays a steady, ongoing stream of validation RPCs from every active client — that per-client validation traffic scales roughly linearly (or worse, under bursty access patterns) with the number of clients and directly consumes server CPU and network capacity regardless of whether any file actually changed. AFS instead uses whole-file caching with server-issued callbacks: once a client has cached a file, it can use it indefinitely without contacting the server again, and the server only needs to send a notification on the (relatively rare) event that the file actually changes. This shifts the server's steady-state load from "field a continuous stream of polling requests from every client" to "track a callback per (client, cached file) pair and rarely send a notification," which is far cheaper per client in the common case where most files are read far more often than they are written — letting one AFS server sustain a substantially larger population of clients than an equivalent NFS server before becoming CPU- or network-bound.

(c) What NFS trades off for a stateless server. A stateless server keeps no record of which files any client currently has open, which is exactly what lets it reboot (or fail over) with no session state to lose or recover — a client simply resends its next request and the server, needing nothing remembered from before, serves it as if nothing happened. This simplicity is bought at the cost of several things a stateful server could otherwise provide cheaply. 1. Weaker cache consistency. With no server-side record of who is caching what, the server cannot proactively notify a client when a file it is caching changes (unlike AFS's callbacks); NFS clients must instead poll (re-validate cached attributes periodically), which means a client can transiently see stale data between polls, and picking the polling interval is itself a tradeoff between staleness and validation traffic. 2. No server-enforced file locking / open semantics across reboots. Because the server keeps no state about which processes have a file open, exclusive-access semantics and advisory locks must be layered on as a separate, often less robust protocol (e.g. NLM alongside classic NFS) rather than being a natural property of a stateful open/close pair, and any such lock state is itself lost if the server (or lock manager) crashes and cannot be trivially resumed. 3. All per-open-file state moves to the client. The client-side NFS module must now hold everything a local filesystem's kernel would otherwise track on the server's behalf — the current file position/offset (since every read/write call must carry an explicit offset argument rather than the server remembering one), the file handle, and the open mode — which is extra client-side complexity that a stateful server design would not require.

(d) AFS vs. NFS — scalability. Even with servers added as required, AFS's own scalability has real limits. Its callback bookkeeping grows with the number of (client, cached-file) pairs a server must track, so a server serving a very large, highly dynamic client population accumulates a correspondingly large amount of per-client state to maintain; a server recovering from a crash must re-establish or invalidate every outstanding callback it had promised, which becomes a heavier recovery burden as the client population grows, unlike NFS's stateless server which has nothing to recover. AFS's cell-based administrative partitioning also means very large deployments must be explicitly divided into cells (each with its own servers and administration), which is an organizational rather than a purely technical scaling limit, but a real one — adding capacity is not simply "add another server to the same pool" the way a stateless, share-nothing NFS deployment can more nearly be. In the other direction, NFS scales its server count easily (each server is independent and stateless, so adding one is trivial), but as noted in (b) it scales its per-server client count worse than AFS because of its polling-based consistency traffic — so the two systems' scalability limits sit in different places: AFS's ceiling is server-side callback/recovery bookkeeping per client, NFS's ceiling is per-client validation traffic per server.