25-Comp-B10 Distributed Systems · December 2016
Nivaar worked solution (AI-drafted; not reviewed by a licensed engineer)
Question text not reproduced: the examination questions are © Engineers and Geoscientists BC. Open the official past paper (linked at the top of this page) to read the question, then follow the worked solution below.
(a) How AFS gains control at open/close. AFS intercepts filesystem calls at the vnode layer inside the client's kernel: when a system call (open or close) names a path within the shared AFS namespace, the kernel's vnode dispatch recognizes it as belonging to AFS and redirects it to the local user-level cache manager process, traditionally called Venus, instead of routing it to the ordinary local filesystem code. Venus intercepts only at open and close, not at every read/write: on open, it checks whether a valid, callback-promised whole-file copy already exists in the local cache; if not (or its callback has since been broken), Venus fetches the entire file from the responsible File Server, installs a callback promise for it, and returns a local file descriptor pointing at the cached copy, so every subsequent read/write the process performs is serviced entirely locally with no further server interaction until the file is closed (at which point, if the file was modified, Venus writes the whole updated copy back to the server). This "gain control only at open/close" design is exactly why AFS is described as whole-file caching, in contrast to a system that must involve the server on every individual read or write.
(b) Per-process state held by the NFS client module. The NFS server itself is deliberately stateless: every read/write request it handles is self-contained, naming an explicit file handle and an explicit byte offset, so the server keeps no record from one call to the next of which files any client currently has "open." All of the state a conventional local file system would keep server-side must therefore be held by the client-side NFS module on behalf of each user-level process that has a file open: (1) the current file position (seek offset) for that file descriptor, since the server's calls take an explicit offset argument rather than remembering one, so the client tracks and supplies it, advancing it locally after each operation; (2) the file handle the server issued when the file was first looked up/opened — an opaque identifier (filesystem id + inode-like reference + generation number) presented on every subsequent operation instead of re-resolving the pathname each time; (3) the open mode (read/write/append) and any locally cached attributes or recently read/written data blocks used to reduce round trips; and (4) the mapping from the process's own local file descriptor number back to that file handle and position, so the descriptor keeps working correctly across many separate, otherwise-stateless RPC calls.
(c) How AFS deals with lost callback messages. If a callback-break notification from the server never reaches a client — lost to a network partition, or the client simply being unreachable — the client would otherwise keep trusting a cached copy that is no longer valid, indefinitely. AFS bounds this risk by giving every callback promise a fixed expiration time (a lease, rather than an unconditional guarantee): once that time elapses, Venus is required to revalidate the callback with the server before trusting the cached copy further, whether or not any break notification was ever received, so a lost break message only risks staleness up to the promise's lifetime, never indefinitely. In addition, whenever a client reboots or reconnects after being unreachable for any period, Venus treats every callback it is holding as suspect and revalidates all of them with the relevant servers before relying on any cached copy again — precisely the situation in which a break message sent while the client could not be reached would otherwise be silently lost forever.
(d) AFS vs. NFS — stability and scalability. Stability/consistency model: classic NFS (v2/v3) is essentially stateless and relies on clients periodically re-validating cached data against the server (a time-based, "check on open"/polling consistency), so the server does no bookkeeping of who is caching what but pays a steady stream of validation requests from every active client, and different clients can transiently see stale data between validations. AFS instead uses whole-file caching with server-issued callbacks: when a client caches a file, the server promises to notify that client if the file changes, so the client can trust its cache until told otherwise without repeatedly re-asking. This shifts load from "many clients constantly polling" to "the server does a small amount of bookkeeping and notifies rarely," which is more stable and scalable under many, mostly-read clients, at the cost of the server tracking per-client callback state. AFS scalability limits: even with servers added as required, AFS's callback bookkeeping grows with the number of (client, cached-file) pairs a server must track, and a server recovering from a crash must re-establish or invalidate every outstanding callback it had promised, a heavier recovery burden as the client population grows; AFS's cell-based administrative partitioning also means very large deployments must be explicitly divided into cells, an organizational rather than a purely technical scaling limit. Recent developments: NFSv4 closed much of the original gap by adopting AFS-style delegations (a server-granted, revocable promise similar to a callback) plus a stronger, compound-RPC, more stateful protocol design, and is now the more actively maintained standard; Coda (built directly on AFS's model) added disconnected operation, letting clients keep working from cache during a total server/network outage and reconcile changes afterward, addressing availability rather than raw throughput scalability.