server-skynet-source-3rd-jemalloc

project-base/server-skynet-source-3rd-jemalloc

Author	SHA1	Message	Date
David Goldblatt	da63f23e68	HPA: Track pending purges/hugifies in the psset. This finishes the refactoring of the HPA/psset interactions the past few commits have been building towards. Rather than the HPA removing and then reinserting hpdatas, it simply begins updates and ends them. These updates can set flags on the hpdata that prevent it from being returned for certain types of requests. For example, it can call hpdata_alloc_allowed_set(hpdata, false) during an update, at which point the given hpdata will no longer be returned for psset_pick_alloc requests. This has various of benefits: - It maintains stats correctness during purges and hugifies. - It allows simpler and more explicit concurrency control for the various special cases (e.g. allocations are disallowed during purge, but not during hugify). - It lets allocations and deallocations avoid disturbing the purging and hugification orderings. If an hpdata "loses its place" in one of the queues just do to an alloc / dalloc, it can result in pathological edge cases where very hot, very full hugepages never get hugified (and cold extents on the same hugepage as hot ones never get purged). The key benefit though is that tracking hpdatas to be purged / hugified in a principled way will let us do delayed purging and hugification. Eventually this will let us move these operations to background threads, but in the short term the benefit is that it will let us have global purging policies (e.g. purge when the entire arena has too many dirty pages, rather than any particular hugepage).	2021-02-04 20:58:31 -08:00
David Goldblatt	0ea3d6307c	CTL, Stats: report HPA empty slab stats.	2021-02-04 20:58:31 -08:00
David Goldblatt	bf64557ed6	Move empty slab tracking to the psset. We're moving towards a world in which purging decisions are less rigidly enforced at a single-hugepage level. In that world, it makes sense to keep around some hpdatas which are not completely purged, in which case we'll need to track them.	2021-02-04 20:58:31 -08:00
David Goldblatt	99fc0717e6	psset: Reconceptualize insertion/removal. Really, this isn't a functional change, just a naming change. We start thinking of pageslabs as being always in the psset. What we used to think of as removal is now thought of as being in the psset, but in the process of being updated (and therefore, unavalable for serving new allocations). This is in preparation of subsequent changes to support deferred purging; allocations will still be in the psset for the purposes of choosing when to purge, but not for purposes of allocation/deallocation.	2021-02-04 20:58:31 -08:00
David Goldblatt	061cabb712	HPA stats: report retained instead of inactive. This more closely maps to the PAC.	2021-02-04 20:58:31 -08:00
David Goldblatt	d3e5ea03c5	HPA: Track dirty stats.	2021-02-04 20:58:31 -08:00
David Goldblatt	68a1666e91	hpdata: Rename "dirty" to "touched". This matches the usage in the rest of the codebase.	2021-02-04 20:58:31 -08:00
David Goldblatt	be0d7a53f3	HPA: Don't track inactive pages. This is really only useful for human consumption. Correspondingly, emit it only in the human-readable stats, and let everybody else compute from the hugepage size and nactive.	2021-02-04 20:58:31 -08:00
David Goldblatt	55e0f60ca1	psset stats: Simplify handling. We can treat the huge and nonhuge cases uniformly using huge state as an array index.	2021-02-04 20:58:31 -08:00
David Goldblatt	94cd9444c5	HPA: Some minor reformattings.	2021-02-04 20:58:31 -08:00
David Goldblatt	b25ee5d88e	HPA: Add purge stats.	2021-02-04 20:58:31 -08:00
David Goldblatt	746ea3de6f	HPA stats: Allow some derived stats. However, we put them in their own struct, to avoid the messiness that the arena has (mixing derived and non-derived stats in the arena_stats_t).	2021-02-04 20:58:31 -08:00
David Goldblatt	30b9e8162b	HPA: Generalize purging. Previously, we would purge a hugepage only when it's completely empty. With this change, we can purge even when only partially empty. Although the heuristic here is still fairly primitive, this infrastructure can scale to become more advanced.	2021-02-04 20:58:31 -08:00
David Goldblatt	70692cfb13	hpdata: Add state changing helpers. We're about to allow hugepage subextent purging; get as much of our metadata handling ready as possible.	2021-02-04 20:58:31 -08:00
David Goldblatt	9b75808be1	flat bitmap: Add a bitwise and/or/not. We're about to need them.	2021-02-04 20:58:31 -08:00
David Goldblatt	2ae966222f	hpdata: track per-page dirty state.	2021-02-04 20:58:31 -08:00
David Goldblatt	ff4086aa6b	hpdata: count active pages instead of free ones. This will be more consistent with later naming choices.	2021-02-04 20:58:31 -08:00
David Goldblatt	3624dd42ff	hpdata: Add a comment for hpdata_consistent.	2021-02-04 20:58:31 -08:00
David Goldblatt	20140629b4	Bin: Move stats closer to the mutex. This is a slight cache locality optimization.	2021-02-04 14:10:43 -08:00
David Goldblatt	c259323ab3	Use ticker_geom_t for arena tcache decay.	2021-02-04 14:10:43 -08:00
David Goldblatt	8edfc5b170	Add ticker_geom_t. This lets a single ticker object drive events across a large number of different tick streams while sharing state.	2021-02-04 14:10:43 -08:00
David Goldblatt	3967329813	Arena: share bin offsets in a global. This saves us a cache miss when lookup up the arena bin offset in a remote arena during tcache flush. All arenas share the base offset, and so we don't need to look it up repeatedly for each arena. Secondarily, it shaves 288 bytes off the arena on, e.g., x86-64.	2021-02-04 14:10:43 -08:00
David Goldblatt	2fcbd18115	Cache bin: Don't reverse flush order. The items we pick to flush matter a lot, but the order in which they get flushed doesn't; just use forward scans. This simplifies the accessing code, both in terms of the C and the generated assembly (i.e. this speeds up the flush pathways).	2021-02-04 14:10:43 -08:00
David Goldblatt	4c46e11365	Cache an arena's index in the arena. This saves us a pointer hop down some perf-sensitive paths.	2021-02-04 14:10:43 -08:00
David Goldblatt	229994a204	Tcache flush: keep common path state in registers. By carefully force-inlining the division constants and the operation sum count, we can eliminate redundant operations in the arena-level dalloc function. Do so.	2021-02-04 14:10:43 -08:00
David Goldblatt	31a629c3de	Tcache flush: prefetch edata contents. This frontloads more of the miss latency. It also moves it to a pathway where we have not yet acquired any locks, so that it should (hopefully) reduce hold times.	2021-02-04 14:10:43 -08:00
David Goldblatt	9f9247a62e	Tcache fluhing: increase cache miss parallelism. In practice, many rtree_leaf_elm accesses are cache misses. By restructuring, we can make it more likely that these misses occur without blocking us from starting later lookups, taking more of those misses in parallel.	2021-02-04 14:10:43 -08:00
David Goldblatt	181ba7fd4d	Tcache flush: Add an emap "batch lookup" path. For now this is a no-op; but the interface is a little more flexible for our purposes.	2021-02-04 14:10:43 -08:00
David Goldblatt	c007c537ff	Tcache flush: Unify edata lookup path.	2021-02-04 14:10:43 -08:00
David CARLIER	35a8552605	Mac OS: Tag mapped pages. This can be used to help profiling tools (e.g. vmmap) identify the sources of mappings more specifically.	2021-02-03 15:05:53 -08:00
Yinan Zhang	f6699803e2	Fix duration in prof log	2021-01-25 16:38:38 -08:00
Azat Khuzhin	a943172b73	Add runtime detection for MADV_DONTNEED zeroes pages (mostly for qemu) qemu does not support this, yet [1], and you can get very tricky assert if you will run program with jemalloc in use under qemu: <jemalloc>: ../contrib/jemalloc/src/extent.c:1195: Failed assertion: "p[i] == 0" [1]: https://patchwork.kernel.org/patch/10576637/ Here is a simple example that shows the problem [2]: // Gist to check possible issues with MADV_DONTNEED // For example it does not supported by qemu user // There is a patch for this [1], but it hasn't been applied. // [1]: https://lists.gnu.org/archive/html/qemu-devel/2018-08/msg05422.html #include <sys/mman.h> #include <stdio.h> #include <stddef.h> #include <assert.h> #include <string.h> int main(int argc, char *argv) { void addr = mmap(NULL, 1<<16, PROT_READ\|PROT_WRITE, MAP_PRIVATE\|MAP_ANONYMOUS, -1, 0); if (addr == MAP_FAILED) { perror("mmap"); return 1; } memset(addr, 'A', 1<<16); if (!madvise(addr, 1<<16, MADV_DONTNEED)) { puts("MADV_DONTNEED does not return error. Check memory."); for (int i = 0; i < 1<<16; ++i) { assert(((unsigned char )addr)[i] == 0); } } else { perror("madvise"); } if (munmap(addr, 1<<16)) { perror("munmap"); return 1; } return 0; } ### unpatched qemu $ qemu-x86_64-static /tmp/test-MADV_DONTNEED MADV_DONTNEED does not return error. Check memory. test-MADV_DONTNEED: /tmp/test-MADV_DONTNEED.c:19: main: Assertion `((unsigned char )addr)[i] == 0' failed. qemu: uncaught target signal 6 (Aborted) - core dumped Aborted (core dumped) ### patched qemu (by returning ENOSYS error) $ qemu-x86_64 /tmp/test-MADV_DONTNEED madvise: Success ### patch for qemu to return ENOSYS diff --git a/linux-user/syscall.c b/linux-user/syscall.c index 897d20c076..5540792e0e 100644 --- a/linux-user/syscall.c +++ b/linux-user/syscall.c @@ -11775,7 +11775,7 @@ static abi_long do_syscall1(void cpu_env, int num, abi_long arg1, turns private file-backed mappings into anonymous mappings. This will break MADV_DONTNEED. This is a hint, so ignoring and returning success is ok. / - return 0; + return ENOSYS; #endif #ifdef TARGET_NR_fcntl64 case TARGET_NR_fcntl64: [2]: https://gist.github.com/azat/12ba2c825b710653ece34dba7f926ece v2: - review fixes - add opt_dont_trust_madvise v3: - review fixes - rename opt_dont_trust_madvise to opt_trust_madvise	2021-01-20 20:08:30 -08:00
Uwe L. Korn	2e3104ba07	Update config.{sub,guess} to support support-aarch64-apple-darwin as a target	2021-01-11 14:00:03 -08:00
David Goldblatt	a011c4c22d	cache_bin: Separate out local and remote accesses. This fixes an incorrect debug-mode assert: - T1 starts an arena stats update and reads stack_head from another thread's cache bin, when that cache bin has 1 item in it. - T2 allocates from that cache bin. The cache_bin's stack_head now points to a NULL pointer, since the cache bin is empty. - T1 Re-reads the cache_bin's stack_head to perform an assertion check (since it previously saw that the bin was empty, whatever stack_head points to should be non-NULL).	2021-01-08 14:18:08 -08:00
Yinan Zhang	14d689c0f9	Add prof stats mutex stats	2021-01-07 20:39:49 -08:00
Yinan Zhang	9f71b5779b	Output prof stats in stats print	2021-01-07 20:39:49 -08:00
Yinan Zhang	1f1a0231ed	Split macros for initializing stats headers	2021-01-07 20:39:49 -08:00
Yinan Zhang	4352cbc21c	Add alignment tests for prof stats	2021-01-07 20:39:49 -08:00
Yinan Zhang	54f3351f1f	Add mallctl for prof stats fetching	2021-01-07 20:39:49 -08:00
Yinan Zhang	40fa4d29d3	Track per size class internal fragmentation	2021-01-07 20:39:49 -08:00
Yinan Zhang	afa489c3c5	Record request size in prof info	2021-01-07 20:39:49 -08:00
David Goldblatt	f9bb8dedef	Un-force-inline do_rallocx. The additional overhead of the function-call setup and flags checking is relatively small, but costs us the replication of the entire realloc pathway in terms of size.	2021-01-04 14:55:49 -08:00
David Goldblatt	a9fa2defdb	Add JEMALLOC_COLD, and mark some functions cold. This hints to the compiler that it should care more about space than CPU (among other things). In cases where the compiler lacks profile-guided information, this can be a substantial space savings. For now, we mark the mallctl or atexit driven profiling and stats functions that take up the most space.	2021-01-04 14:55:49 -08:00
David Goldblatt	5d8e70ab26	prof_recent: cassert(config_prof) more often. This tells the compiler that these functions are never called, which lets them be optimized away in builds where profiling is disabled.	2021-01-04 14:55:49 -08:00
David Goldblatt	83cad746ae	prof_log: cassert(config_prof) in public functions This lets the compiler infer that the code is dead in builds where profiling is enabled, saving on space there.	2021-01-04 14:55:49 -08:00
David Goldblatt	526180b76d	Extent.c: Avoid an rtree NULL-check. The edge case in which pages_map returns (void *)PAGE can trigger an incorrect assertion failure. Avoid it.	2021-01-04 14:50:49 -08:00
Yinan Zhang	b35ac00d58	Do not bump to large size for page aligned request	2020-12-29 17:09:58 -08:00
Yinan Zhang	8a56d6b636	Add last-N mutex stats	2020-12-29 09:44:19 -08:00
Yinan Zhang	22d62d8cbd	Handle ending gap properly for HPA stats	2020-12-18 16:40:57 -08:00
Yinan Zhang	6c5a3a24dd	Omit bin stats rows with no data	2020-12-18 16:40:57 -08:00

1 2 3 4 5 ...

3127 Commits