提交 · 94206125c4aac32e43c25bfe1b827e7ab993b7dc · openeuler / raspberrypi-kernel

12 7月, 2012 2 次提交

ipv4: Rearrange arguments to ip_rt_redirect() · 94206125

由 David S. Miller 提交于 7月 11, 2012

Pass in the SKB rather than just the IP addresses, so that policy
and other aspects can reside in ip_rt_redirect() rather then
icmp_redirect().
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

94206125

tcp: TCP Small Queues · 46d3ceab

由 Eric Dumazet 提交于 7月 11, 2012

This introduce TSQ (TCP Small Queues)

TSQ goal is to reduce number of TCP packets in xmit queues (qdisc &
device queues), to reduce RTT and cwnd bias, part of the bufferbloat
problem.

sk->sk_wmem_alloc not allowed to grow above a given limit,
allowing no more than ~128KB [1] per tcp socket in qdisc/dev layers at a
given time.

TSO packets are sized/capped to half the limit, so that we have two
TSO packets in flight, allowing better bandwidth use.

As a side effect, setting the limit to 40000 automatically reduces the
standard gso max limit (65536) to 40000/2 : It can help to reduce
latencies of high prio packets, having smaller TSO packets.

This means we divert sock_wfree() to a tcp_wfree() handler, to
queue/send following frames when skb_orphan() [2] is called for the
already queued skbs.

Results on my dev machines (tg3/ixgbe nics) are really impressive,
using standard pfifo_fast, and with or without TSO/GSO.

Without reduction of nominal bandwidth, we have reduction of buffering
per bulk sender :
< 1ms on Gbit (instead of 50ms with TSO)
< 8ms on 100Mbit (instead of 132 ms)

I no longer have 4 MBytes backlogged in qdisc by a single netperf
session, and both side socket autotuning no longer use 4 Mbytes.

As skb destructor cannot restart xmit itself ( as qdisc lock might be
taken at this point ), we delegate the work to a tasklet. We use one
tasklest per cpu for performance reasons.

If tasklet finds a socket owned by the user, it sets TSQ_OWNED flag.
This flag is tested in a new protocol method called from release_sock(),
to eventually send new segments.

[1] New /proc/sys/net/ipv4/tcp_limit_output_bytes tunable
[2] skb_orphan() is usually called at TX completion time,
  but some drivers call it in their start_xmit() handler.
  These drivers should at least use BQL, or else a single TCP
  session can still fill the whole NIC TX ring, since TSQ will
  have no effect.
Signed-off-by: NEric Dumazet <edumazet@google.com>
Cc: Dave Taht <dave.taht@bufferbloat.net>
Cc: Tom Herbert <therbert@google.com>
Cc: Matt Mathis <mattmathis@google.com>
Cc: Yuchung Cheng <ycheng@google.com>
Cc: Nandita Dukkipati <nanditad@google.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

46d3ceab

11 7月, 2012 11 次提交

ipv6: optimize ipv6 addresses compares · 1a203cb3

由 Eric Dumazet 提交于 7月 10, 2012

On 64 bit arches having efficient unaligned accesses (eg x86_64) we can
use long words to reduce number of instructions for free.

Joe Perches suggested to change ipv6_masked_addr_cmp() to return a bool
instead of 'int', to make sure ipv6_masked_addr_cmp() cannot be used
in a sorting function.
Signed-off-by: NEric Dumazet <edumazet@google.com>
Cc: Joe Perches <joe@perches.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

1a203cb3

D
ipv4: Remove inetpeer from routes. · f185071d
由 David S. Miller 提交于 7月 10, 2012
```
No longer used.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
f185071d

ipv4: Maintain redirect and PMTU info in struct rtable again. · 5943634f

由 David S. Miller 提交于 7月 10, 2012

Maintaining this in the inetpeer entries was not the right way to do
this at all.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

5943634f

inet: Kill FLOWI_FLAG_PRECOW_METRICS. · 3e12939a

由 David S. Miller 提交于 7月 10, 2012

No longer needed.  TCP writes metrics, but now in it's own special
cache that does not dirty the route metrics.  Therefore there is no
longer any reason to pre-cow metrics in this way.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

3e12939a

D
inet: Remove ->get_peer() method. · 16d18399
由 David S. Miller 提交于 7月 10, 2012
```
No longer used.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
16d18399
D
tcp: Move timestamps from inetpeer to metrics cache. · 81166dd6
由 David S. Miller 提交于 7月 10, 2012
```
With help from Lin Ming.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
81166dd6
D
net: Kill set_dst_metric_rtt(). · 94334d5e
由 David S. Miller 提交于 7月 10, 2012
```
No longer used.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
94334d5e

tcp: Maintain dynamic metrics in local cache. · 51c5d0c4

由 David S. Miller 提交于 7月 10, 2012

Maintain a local hash table of TCP dynamic metrics blobs.

Computed TCP metrics are no longer maintained in the route metrics.

The table uses RCU and an extremely simple hash so that it has low
latency and low overhead.  A simple hash is legitimate because we only
make metrics blobs for fully established connections.

Some tweaking of the default hash table sizes, metric timeouts, and
the hash chain length limit certainly could use some tweaking.  But
the basic design seems sound.

With help from Eric Dumazet and Joe Perches.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

51c5d0c4

D
tcp: Abstract back handling peer aliveness test into helper function. · ab92bb2f
由 David S. Miller 提交于 7月 09, 2012
```
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
ab92bb2f
D
tcp: Move dynamnic metrics handling into seperate file. · 4aabd8ef
由 David S. Miller 提交于 7月 09, 2012
```
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
4aabd8ef

ipv4: Fix crashes in fib_rules_tclass(). · e044a651

由 David S. Miller 提交于 7月 10, 2012

All paths assume, when CONFIG_IP_MULTIPLE_TABLES is enabled, that any
successful call to fib_lookup() will initialize the fib_result->r
value to something.

We violated that expectation in the new fib_lookup() fast path.
Reported-by: NOr Gerlitz <ogerlitz@mellanox.com>
Tested-by: NEric Dumazet <eric.dumazet@gmail.com>
Tested-by: NGreg Rose <gregory.v.rose@intel.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

e044a651

09 7月, 2012 1 次提交

netfilter: nf_ct_ecache: fix crash with multiple containers, one shutting down · 6bd0405b

由 Pablo Neira Ayuso 提交于 7月 05, 2012

Hans reports that he's still hitting:

BUG: unable to handle kernel NULL pointer dereference at 000000000000027c
IP: [<ffffffff813615db>] netlink_has_listeners+0xb/0x60
PGD 0
Oops: 0000 [#3] PREEMPT SMP
CPU 0

It happens when adding a number of containers with do:

nfct_query(h, NFCT_Q_CREATE, ct);

and most likely one namespace shuts down.

this problem was supposed to be fixed by:
70e9942f netfilter: nf_conntrack: make event callback registration per-netns

Still, it was missing one rcu_access_pointer to check if the callback
is set or not.
Reported-by: NHans Schillstrom <hans@schillstrom.com>
Signed-off-by: NPablo Neira Ayuso <pablo@netfilter.org>

6bd0405b

06 7月, 2012 1 次提交

ipv4: Avoid overhead when no custom FIB rules are installed. · f4530fa5

由 David S. Miller 提交于 7月 05, 2012

If the user hasn't actually installed any custom rules, or fiddled
with the default ones, don't go through the whole FIB rules layer.

It's just pure overhead.

Instead do what we do with CONFIG_IP_MULTIPLE_TABLES disabled, check
the individual tables by hand, one by one.

Also, move fib_num_tclassid_users into the ipv4 network namespace.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

f4530fa5

05 7月, 2012 8 次提交

D
net: Kill dst->_neighbour, accessors, and final uses. · 36bdbcae
由 David S. Miller 提交于 7月 02, 2012
```
No longer used.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
36bdbcae

ipv6: Store route neighbour in rt6_info struct. · 97cac082

由 David S. Miller 提交于 7月 02, 2012

This makes for a simplified conversion away from dst_get_neighbour*().

All code outside of ipv6 will use neigh lookups via dst_neigh_lookup*().
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

97cac082

D
net: Pass neighbours and dest address into NETEVENT_REDIRECT events. · 1d248b1c
由 David S. Miller 提交于 7月 03, 2012
```
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
1d248b1c

decnet: Use neighbours privately in dn_route struct. · fccd7d5c

由 David S. Miller 提交于 7月 02, 2012

This allows an easy conversion away from dst_get_neighbour*().
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

fccd7d5c

net: Add optional SKB arg to dst_ops->neigh_lookup(). · f894cbf8

由 David S. Miller 提交于 7月 02, 2012

Causes the handler to use the daddr in the ipv4/ipv6 header when
the route gateway is unspecified (local subnet).
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

f894cbf8

net: Do delayed neigh confirmation. · 5110effe

由 David S. Miller 提交于 7月 02, 2012

When a dst_confirm() happens, mark the confirmation as pending in the
dst.  Then on the next packet out, when we have the neigh in-hand, do
the update.

This removes the dependency in dst_confirm() of dst's having an
attached neigh.

While we're here, remove the explicit 'dst' NULL check, all except 2
or 3 call sites ensure it's not NULL.  So just fix those cases up.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

5110effe

ipv4: Make neigh lookups directly in output packet path. · a263b309

由 David S. Miller 提交于 7月 02, 2012

Do not use the dst cached neigh, we'll be getting rid of that.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

a263b309

netfilter: nf_conntrack: generalize nf_ct_l4proto_net · 08911475

由 Pablo Neira Ayuso 提交于 6月 29, 2012

This patch generalizes nf_ct_l4proto_net by splitting it into chunks and
moving the corresponding protocol part to where it really belongs to.

To clarify, note that we follow two different approaches to support per-net
depending if it's built-in or run-time loadable protocol tracker.
Signed-off-by: NPablo Neira Ayuso <pablo@netfilter.org>
Acked-by: NGao feng <gaofeng@cn.fujitsu.com>

08911475

01 7月, 2012 1 次提交

sctp: be more restrictive in transport selection on bundled sacks · 4244854d

由 Neil Horman 提交于 6月 30, 2012

It was noticed recently that when we send data on a transport, its possible that
we might bundle a sack that arrived on a different transport. While this isn't
a major problem, it does go against the SHOULD requirement in section 6.4 of RFC
2960:

An endpoint SHOULD transmit reply chunks (e.g., SACK, HEARTBEAT ACK,
etc.) to the same destination transport address from which it
received the DATA or control chunk to which it is replying. This
rule should also be followed if the endpoint is bundling DATA chunks
together with the reply chunk.

This patch seeks to correct that. It restricts the bundling of sack operations
to only those transports which have moved the ctsn of the association forward
since the last sack. By doing this we guarantee that we only bundle outbound
saks on a transport that has received a chunk since the last sack. This brings
us into stricter compliance with the RFC.

Vlad had initially suggested that we strictly allow only sack bundling on the
transport that last moved the ctsn forward. While this makes sense, I was
concerned that doing so prevented us from bundling in the case where we had
received chunks that moved the ctsn on multiple transports. In those cases, the
RFC allows us to select any of the transports having received chunks to bundle
the sack on. so I've modified the approach to allow for that, by adding a state
variable to each transport that tracks weather it has moved the ctsn since the
last sack. This I think keeps our behavior (and performance), close enough to
our current profile that I think we can do this without a sysctl knob to
enable/disable it.
Signed-off-by: NNeil Horman <nhorman@tuxdriver.com>
CC: Vlad Yaseivch <vyasevich@gmail.com>
CC: David S. Miller <davem@davemloft.net>
CC: linux-sctp@vger.kernel.org
Reported-by: NMichele Baldessari <michele@redhat.com>
Reported-by: Nsorin serban <sserban@redhat.com>
Acked-by: NVlad Yasevich <vyasevich@gmail.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

4244854d

29 6月, 2012 5 次提交

ipv4: Elide fib_validate_source() completely when possible. · 7a9bc9b8

由 David S. Miller 提交于 6月 29, 2012

If rpfilter is off (or the SKB has an IPSEC path) and there are not
tclassid users, we don't have to do anything at all when
fib_validate_source() is invoked besides setting the itag to zero.

We monitor tclassid uses with a counter (modified only under RTNL and
marked __read_mostly) and we protect the fib_validate_source() real
work with a test against this counter and whether rpfilter is to be
done.

Having a way to know whether we need no tclassid processing or not
also opens the door for future optimized rpfilter algorithms that do
not perform full FIB lookups.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

7a9bc9b8

ipv6_tunnel: Allow receiving packets on the fallback tunnel if they pass sanity checks · d0087b29

由 Ville Nuorvala 提交于 6月 28, 2012

At Facebook, we do Layer-3 DSR via IP-in-IP tunneling. Our load balancers wrap
an extra IP header on incoming packets so they can be routed to the backend.
In the v4 tunnel driver, when these packets fall on the default tunl0 device,
the behavior is to decapsulate them and drop them back on the stack. So our
setup is that tunl0 has the VIP and eth0 has (obviously) the backend's real
address.

In IPv6 we do the same thing, but the v6 tunnel driver didn't have this same
behavior - if you didn't have an explicit tunnel setup, it would drop the
packet.

This patch brings that v4 feature to the v6 driver.

The same IPv6 address checks are performed as with any normal tunnel,
but as the fallback tunnel endpoint addresses are unspecified, the checks
must be performed on a per-packet basis, rather than at tunnel
configuration time.

[Patch description modified by phil@ipom.com]
Signed-off-by: NVille Nuorvala <ville.nuorvala@gmail.com>
Tested-by: NPhil Dibowitz <phil@ipom.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

d0087b29

ipv4: Adjust in_dev handling in fib_validate_source() · 9e56e380

由 David S. Miller 提交于 6月 28, 2012

Checking for in_dev being NULL is pointless.

In fact, all of our callers have in_dev precomputed already,
so just pass it in and remove the NULL checking.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

9e56e380

net: Use NLMSG_DEFAULT_SIZE in combination with nlmsg_new() · 58050fce

由 Thomas Graf 提交于 6月 28, 2012

Using NLMSG_GOODSIZE results in multiple pages being used as
nlmsg_new() will automatically add the size of the netlink
header to the payload thus exceeding the page limit.

NLMSG_DEFAULT_SIZE takes this into account.
Signed-off-by: NThomas Graf <tgraf@suug.ch>
Cc: Jiri Pirko <jpirko@redhat.com>
Cc: Dmitry Eremin-Solenikov <dbaryshkov@gmail.com>
Cc: Sergey Lapin <slapin@ossfans.org>
Cc: Johannes Berg <johannes@sipsolutions.net>
Cc: Lauro Ramos Venancio <lauro.venancio@openbossa.org>
Cc: Aloisio Almeida Jr <aloisio.almeida@openbossa.org>
Cc: Samuel Ortiz <sameo@linux.intel.com>
Reviewed-by: NJiri Pirko <jpirko@redhat.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

58050fce

tcp: pass fl6 to inet6_csk_route_req() · 3840a06e

由 Neal Cardwell 提交于 6月 28, 2012

This commit changes inet_csk_route_req() so that it uses a pointer to
a struct flowi6, rather than allocating its own on the stack. This
brings its behavior in line with its IPv4 cousin,
inet_csk_route_req(), and allows a follow-on patch to fix a dst leak.
Signed-off-by: NNeal Cardwell <ncardwell@google.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

3840a06e

28 6月, 2012 9 次提交

D
ipv4: Kill rt->rt_spec_dst, no longer used. · 41347dcd
由 David S. Miller 提交于 6月 28, 2012
```
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
41347dcd

ipv4: Create and use fib_compute_spec_dst() helper. · 35ebf65e

由 David S. Miller 提交于 6月 28, 2012

The specific destination is the host we direct unicast replies to.
Usually this is the original packet source address, but if we are
responding to a multicast or broadcast packet we have to use something
different.

Specifically we must use the source address we would use if we were to
send a packet to the unicast source of the original packet.

The routing cache precomputes this value, but we want to remove that
precomputation because it creates a hard dependency on the expensive
rpfilter source address validation which we'd like to make cheaper.

There are only three places where this matters:

1) ICMP replies.

2) pktinfo CMSG

3) IP options

Now there will be no real users of rt->rt_spec_dst and we can simply
remove it altogether.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

35ebf65e

ipv4: Show that ip_send_reply() is purely unicast routine. · 70e73416

由 David S. Miller 提交于 6月 28, 2012

Rename it to ip_send_unicast_reply() and add explicit 'saddr'
argument.

This removed one of the few users of rt->rt_spec_dst.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

70e73416

D
ipv4: Kill early demux method return value. · 160eb5a6
由 David S. Miller 提交于 6月 27, 2012
```
It's completely unnecessary.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>
```
160eb5a6

xfrm_user: Propagate netlink error codes properly. · 1d1e34dd

由 David S. Miller 提交于 6月 27, 2012

Instead of using a fixed value of "-1" or "-EMSGSIZE", propagate what
the nla_*() interfaces actually return.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

1d1e34dd

Revert "ipv4: tcp: dont cache unconfirmed intput dst" · c10237e0

由 David S. Miller 提交于 6月 27, 2012

This reverts commit c074da28.

This change has several unwanted side effects:

1) Sockets will cache the DST_NOCACHE route in sk->sk_rx_dst and we'll
   thus never create a real cached route.

2) All TCP traffic will use DST_NOCACHE and never use the routing
   cache at all.
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

c10237e0

ipv4: tcp: dont cache unconfirmed intput dst · c074da28

由 Eric Dumazet 提交于 6月 26, 2012

DDOS synflood attacks hit badly IP route cache.

On typical machines, this cache is allowed to hold up to 8 Millions dst
entries, 256 bytes for each, for a total of 2GB of memory.

rt_garbage_collect() triggers and tries to cleanup things.

Eventually route cache is disabled but machine is under fire and might
OOM and crash.

This patch exploits the new TCP early demux, to set a nocache
boolean in case incoming TCP frame is for a not yet ESTABLISHED or
TIMEWAIT socket.

This 'nocache' boolean is then used in case dst entry is not found in
route cache, to create an unhashed dst entry (DST_NOCACHE)

SYN-cookie-ACK sent use a similar mechanism (ipv4: tcp: dont cache
output dst for syncookies), so after this patch, a machine is able to
absorb a DDOS synflood attack without polluting its IP route cache.
Signed-off-by: NEric Dumazet <edumazet@google.com>
Cc: Hans Schillstrom <hans.schillstrom@ericsson.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

c074da28

netfilter: nf_conntrack: add nf_ct_kfree_compat_sysctl_table · f28997e2

由 Gao feng 提交于 6月 21, 2012

This patch is a cleanup.

It adds nf_ct_kfree_compat_sysctl_table to release l4proto's
compat sysctl table and set the compat sysctl table point to NULL.

This new function will be used by follow-up patches.
Signed-off-by: NGao feng <gaofeng@cn.fujitsu.com>
Signed-off-by: NPablo Neira Ayuso <pablo@netfilter.org>

f28997e2

netfilter: nf_conntrack: prepare l4proto->init_net cleanup · f1caad27

由 Gao feng 提交于 6月 21, 2012

l4proto->init contain quite redundant code. We can simplify this
by adding a new parameter l3proto.

This patch prepares that code simplification.
Signed-off-by: NGao feng <gaofeng@cn.fujitsu.com>
Signed-off-by: NPablo Neira Ayuso <pablo@netfilter.org>

f1caad27

27 6月, 2012 1 次提交

mac802154: add wpan device-class support · 32bad7e3

由 alex.bluesman.smirnov@gmail.com 提交于 6月 25, 2012

Every real 802.15.4 transceiver, which works with software MAC layer,
can be classified as a wpan device in this stack. So the wpan device
implementation provides missing link in datapath between the device
drivers and the Linux network queue.

According to the IEEE 802.15.4 standard each packet can be one of the
following types:
 - beacon
 - MAC layer command
 - ACK
 - data

This patch adds support for the data packet-type only, but this is
enough to perform data transmission and receiving over radio.
Signed-off-by: NAlexander Smirnov <alex.bluesman.smirnov@gmail.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

32bad7e3

26 6月, 2012 1 次提交

nl80211: specify RSSI threshold in scheduled scan · 88e920b4

由 Thomas Pedersen 提交于 6月 21, 2012

Support configuring an RSSI threshold in dBm (s32) when requesting
scheduled scan, below which a BSS won't be reported by the cfg80211
driver.
Signed-off-by: NThomas Pedersen <c_tpeder@qca.qualcomm.com>
Signed-off-by: NJohannes Berg <johannes.berg@intel.com>

88e920b4